@agentproto/apps 0.18.0 → 0.20.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (32) hide show
  1. package/dist/index.d.ts +2 -0
  2. package/dist/index.d.ts.map +1 -1
  3. package/dist/index.mjs +61 -2
  4. package/dist/index.mjs.map +1 -1
  5. package/dist/review-panel/panel.d.ts +1 -1
  6. package/dist/review-panel/panel.d.ts.map +1 -1
  7. package/dist/review-panel/panel.generated.d.ts +1 -1
  8. package/dist/review-panel/panel.generated.d.ts.map +1 -1
  9. package/dist/review-panel/panel.mjs +1 -1
  10. package/dist/review-panel/panel.mjs.map +1 -1
  11. package/dist/review-panel.mjs +1 -1
  12. package/dist/review-panel.mjs.map +1 -1
  13. package/dist/store/index.d.ts +61 -0
  14. package/dist/store/index.d.ts.map +1 -0
  15. package/dist/store/panel.d.ts +53 -0
  16. package/dist/store/panel.d.ts.map +1 -0
  17. package/dist/store/panel.generated.d.ts +13 -0
  18. package/dist/store/panel.generated.d.ts.map +1 -0
  19. package/dist/store/panel.mjs +25 -0
  20. package/dist/store/panel.mjs.map +1 -0
  21. package/dist/store.mjs +71 -0
  22. package/dist/store.mjs.map +1 -0
  23. package/package.json +14 -2
  24. package/session-steward/.agentproto/APP.md +2 -0
  25. package/session-steward/.agentproto/workflows/session-steward/WORKFLOW.md +218 -11
  26. package/session-steward/.agentproto/workflows/session-steward/cron-rules.mjs +526 -0
  27. package/session-steward/.agentproto/workflows/session-steward/entry.mjs +539 -64
  28. package/session-steward/.agentproto/workflows/session-steward/origin-policy.mjs +115 -0
  29. package/session-steward/README.md +19 -2
  30. package/session-steward/routines/session-steward-hourly/ROUTINE.md +24 -1
  31. package/session-steward/scripts/sessions-snapshot.sh +85 -0
  32. package/session-steward/skill/SKILL.md +74 -0
@@ -53,6 +53,43 @@ inputs:
53
53
  callerSessionId:
54
54
  type: string
55
55
  description: The calling session's id — never a candidate.
56
+ callerOrigin:
57
+ type: string
58
+ description: >-
59
+ The calling session's origin (`cron:<jobId>`) — an older run of the
60
+ SAME cron job is never judged as user work.
61
+ appId:
62
+ type: string
63
+ description: >-
64
+ Installed app whose `app_state` ledger holds the verdict memory.
65
+ Default `@agentproto/session-steward` (the id in APP.md). A bare name
66
+ also matches a scoped install; a missing app means no memory, noted in
67
+ the report.
68
+ stableVerdictPasses:
69
+ type: number
70
+ description: >-
71
+ Consecutive passes on an unchanged evidence fingerprint before the
72
+ judge cache stops re-judging a session.
73
+ default: 2
74
+ userOrigins:
75
+ type: array
76
+ description: >-
77
+ Origins that are ALWAYS flag-only, never closed (a human is in the
78
+ loop). A trailing `*` is a prefix wildcard. Default
79
+ `["chat-starter", "vscode"]`; a root with no origin and no parent is
80
+ treated as a user origin too.
81
+ items:
82
+ type: string
83
+ default: ["chat-starter", "vscode"]
84
+ closableOrigins:
85
+ type: array
86
+ description: >-
87
+ Origins that may be closed under the current rules. A trailing `*` is a
88
+ prefix wildcard. Default `["cron:*", "gate"]`. Executors (a session with
89
+ a `parentSessionId`) are closable regardless.
90
+ items:
91
+ type: string
92
+ default: ["cron:*", "gate"]
56
93
  outputs: {}
57
94
  steps:
58
95
  - id: modelRoles
@@ -75,6 +112,7 @@ steps:
75
112
  tool: session_wrapup_plan
76
113
  inputs:
77
114
  idleMinutes: $steps.settings.idleMinutes
115
+ wait: true
78
116
 
79
117
  - id: candidates
80
118
  kind: transform
@@ -83,11 +121,36 @@ steps:
83
121
  Entry-based — splitCandidates. `judge` is ordered most RAM first and
84
122
  capped at `maxJudged`.
85
123
 
124
+ - id: hostLoad
125
+ kind: tool
126
+ name: Host saturation report (report only)
127
+ tool: host_load
128
+ inputs: {}
129
+
130
+ - id: liveSessions
131
+ kind: tool
132
+ name: List live sessions for the mechanical scan
133
+ tool: session_list
134
+ inputs:
135
+ full: true
136
+
137
+ - id: scan
138
+ kind: transform
139
+ name: Scan live sessions (busy / idle / terminal / never-ran / excluded)
140
+ description: Entry-based — scanLive.
141
+
142
+ - id: candidatesPlus
143
+ kind: transform
144
+ name: Merge never-ran 0/0 sessions into stuck
145
+ description: Entry-based — mergeNeverRan (never-ran is stuck, never judged).
146
+
86
147
  - id: ruleApplyQueue
87
148
  kind: transform
88
149
  name: Rule verdicts to apply
89
150
  description: >-
90
151
  Entry-based. Empty unless `apply`: `close` → done, `stuck` → abandoned.
152
+ Origin-bounded — a user-origin candidate is queued as a `needs-input`
153
+ FLAG instead of a close.
91
154
 
92
155
  - id: autoApply
93
156
  kind: map
@@ -104,10 +167,81 @@ steps:
104
167
  verdict: $item.verdict
105
168
  note: $item.note
106
169
 
170
+ - id: installedApps
171
+ kind: tool
172
+ name: List installed apps (to resolve the memory app)
173
+ tool: app_list
174
+ inputs: {}
175
+
176
+ - id: memoryApp
177
+ kind: transform
178
+ name: Resolve the installed app holding the verdict memory
179
+ description: Entry-based — resolveMemoryApp. No such app ⇒ no memory + a report note.
180
+
181
+ - id: memoryQueue
182
+ kind: transform
183
+ name: Memory read queue
184
+ description: Entry-based. Empty when no installed memory app resolves.
185
+
186
+ - id: memoryRead
187
+ kind: map
188
+ name: Read the verdict memory ledger (best-effort)
189
+ over: $steps.memoryQueue
190
+ parallelism: 1
191
+ onError: collect
192
+ steps:
193
+ - id: memoryReadOne
194
+ kind: tool
195
+ tool: app_state_list
196
+ inputs:
197
+ appId: $item.appId
198
+ stage: session-steward
199
+ kinds: [note]
200
+ limit: 500
201
+
202
+ - id: memory
203
+ kind: transform
204
+ name: Fold verdict memory
205
+ description: Entry-based — foldMemory.
206
+
207
+ - id: loopScan
208
+ kind: map
209
+ name: Loop sanity per busy session
210
+ over: $steps.scan.loopQueue
211
+ parallelism: 4
212
+ onError: collect
213
+ steps:
214
+ - id: loopCallsOne
215
+ kind: tool
216
+ tool: tool_calls_list
217
+ inputs:
218
+ sessionId: $item.sessionId
219
+ lastN: 60
220
+ - id: loopFold
221
+ kind: transform
222
+ name: Loop verdict for one session
223
+
224
+ - id: loopResults
225
+ kind: transform
226
+ name: Loop verdicts
227
+ description: Entry-based.
228
+
229
+ - id: proposals
230
+ kind: transform
231
+ name: Nudge proposals (report only — never executed here)
232
+ description: >-
233
+ Entry-based — buildProposalsStep: loop → interrupt, stall → continue, at
234
+ most one per session per pass; a user-origin session is observed only.
235
+
236
+ - id: relabelQueue
237
+ kind: transform
238
+ name: Terminal sessions missing an outcome (relabel candidates)
239
+ description: Entry-based — buildRelabelQueue.
240
+
107
241
  - id: evidence
108
242
  kind: map
109
243
  name: Collect compact evidence per judge candidate
110
- over: $steps.candidates.judge
244
+ over: $steps.candidatesPlus.judge
111
245
  parallelism: 4
112
246
  onError: collect
113
247
  steps:
@@ -119,8 +253,8 @@ steps:
119
253
 
120
254
  - id: judgeQueue
121
255
  kind: transform
122
- name: Candidates with evidence
123
- description: Entry-based.
256
+ name: Candidates with evidence minus the stable-verdict cache
257
+ description: Entry-based — buildJudgeQueueFiltered.
124
258
 
125
259
  - id: jevQueue
126
260
  kind: transform
@@ -207,7 +341,8 @@ steps:
207
341
  name: Confident verdicts to apply
208
342
  description: >-
209
343
  Entry-based. Empty unless `apply`: done/abandoned/blocked/needs-input at
210
- or above `minConfidence`.
344
+ or above `minConfidence`. Origin-bounded — a user-origin candidate is
345
+ downgraded to a `needs-input` FLAG, never a close.
211
346
 
212
347
  - id: judgedApply
213
348
  kind: map
@@ -225,6 +360,25 @@ steps:
225
360
  judgedBy: $item.judgedBy
226
361
  note: $item.note
227
362
 
363
+ - id: memoryWriteQueue
364
+ kind: transform
365
+ name: Verdict memory events to append
366
+ description: Entry-based — buildMemoryWriteQueue (a ledger write, never a session action).
367
+
368
+ - id: memoryWrite
369
+ kind: map
370
+ name: Append verdict memory to app_state (best-effort)
371
+ over: $steps.memoryWriteQueue
372
+ parallelism: 1
373
+ onError: collect
374
+ steps:
375
+ - id: memoryWriteOne
376
+ kind: tool
377
+ tool: app_state_append
378
+ inputs:
379
+ appId: $item.appId
380
+ event: $item.event
381
+
228
382
  - id: report
229
383
  kind: transform
230
384
  name: Build the markdown report
@@ -233,23 +387,67 @@ steps:
233
387
  result:
234
388
  report: $steps.report
235
389
  apply: $steps.settings.apply
236
- candidates: $steps.candidates
390
+ candidates: $steps.candidatesPlus
237
391
  verdicts: $steps.finalVerdicts
238
392
  autoApply: $steps.autoApply
239
393
  judgedApply: $steps.judgedApply
394
+ proposals: $steps.proposals
395
+ relabel: $steps.relabelQueue
396
+ scan: $steps.scan
240
397
  ---
241
398
 
242
399
  # Session Steward — `session-steward` workflow
243
400
 
244
- `session_wrapup_plan` → rules pass over `close`/`stuck` → compact evidence per
245
- `judge` session → one cheap judge turn each → (opt-in) ask the session itself →
246
- close or flag confident verdicts through `session_wrapup_apply` → markdown
247
- report with RAM freed / still held.
401
+ `session_wrapup_plan` → deterministic mechanical pass (host saturation, loop /
402
+ stall / never-ran / terminal-relabel scan, verdict memory) → rules pass over
403
+ `close`/`stuck` → compact evidence per `judge` session → one cheap judge turn
404
+ each (cached when the evidence is unchanged) → (opt-in) ask the session itself
405
+ → close or flag confident verdicts through `session_wrapup_apply` → markdown
406
+ report with proposals, relabel candidates, and RAM freed / still held.
407
+
408
+ ## Mechanical rules (ported from the `kill-idle-sessions` cron prototype)
409
+
410
+ Every rule below is a pure function in `cron-rules.mjs`, pinned by
411
+ `session-steward-cron-rules.test.ts`:
412
+
413
+ - **Loop (1).** `tool_calls_list` per busy session: the same argv verbatim ≥3
414
+ in 10 min, distinct/total < 0.2, or the same file read ≥4 → `looping`, a
415
+ sub-case of `active`. The proposed action is an **interrupt nudge**, never a
416
+ close. Useful loops (watch, test/type-check re-runs, `git status`, `gh pr`
417
+ polling) are excluded.
418
+ - **Stall (2).** Busy > 20 min with no new activity, or a recent
419
+ `lastTurnErroredAt` on an idle process → a proposed **"continue" nudge**.
420
+ - **Never-ran (3).** `tokensIn === 0 && tokensOut === 0` → `stuck`
421
+ immediately, without a judge, whatever the idle.
422
+ - **Fast-path done (4).** Last tool call is `message_parent(kind:done)` plus a
423
+ commit/PR → `done` without a judge (used by the criteria, see below).
424
+ - **Terminal without outcome (5).** Terminal sessions missing an outcome are
425
+ surfaced as relabel candidates instead of staying invisible.
426
+ - **Re-check at apply (6).** A candidate that became busy before the apply is
427
+ skipped (the apply tool also re-classifies).
428
+ - **Self-exclusion (7).** An older run of the caller's own `cron:<job>` is
429
+ never judged as user work.
430
+ - **Explicit 0-candidate report (8).** When nothing is idle, the report says
431
+ why (`n live, m busy, k terminal, j excluded`).
432
+ - **Host saturation (9).** If `host_load` is critical, the report lists
433
+ orphans and big non-session processes FIRST — report only, no action.
434
+ - **Verdict memory (10).** Each verdict is written to the app's `app_state`
435
+ ledger; a session judged the same verdict on an unchanged evidence
436
+ fingerprint for `stableVerdictPasses` passes is served from cache and not
437
+ re-judged. Operator disagreements are recorded as examples.
438
+
439
+ The proposed nudges (interrupt / continue) are **reported only** — this
440
+ workflow never sends a prompt and never closes a `looping` session. At most
441
+ one nudge is proposed per session per pass; a user-origin session is reported
442
+ as observed, never nudged.
248
443
 
249
444
  ## Safety
250
445
 
251
- - `apply: false` (the default) mutates nothing: every mutating map runs over
252
- an empty list.
446
+ - `apply: false` (the default) mutates no SESSION: every session-mutating map
447
+ runs over an empty list. The one write a dry run performs is the append-only
448
+ verdict-memory ledger (`app_state_append`) — it never touches a session and
449
+ is what lets the cache accumulate across passes. Set `appId: ""` to disable
450
+ it entirely.
253
451
  - `session_wrapup_apply` re-classifies each id right before acting and always
254
452
  refuses `keep`-class ids; this workflow never feeds it one.
255
453
  - Rules only ever close `close`/`stuck` ids; a `keepAlive` session is never
@@ -258,6 +456,15 @@ report with RAM freed / still held.
258
456
  - A malformed judge reply is `active` with confidence 0 — never acted on.
259
457
  - The caller's own session (`callerSessionId`) is dropped from every list.
260
458
  - `blocked` / `needs-input` only FLAG a session; it keeps running.
459
+ - **Origin bound (never close a human's session).** Every candidate's
460
+ `origin`/`parentSessionId` runs through the pure `decideAction`
461
+ (`origin-policy.mjs`): a `userOrigins` match (`chat-starter`, `vscode` by
462
+ default) or a root with no origin and no parent is FLAG-ONLY, even with
463
+ `apply: true` and a confident `done` verdict. `cron:*`, `gate`, and
464
+ executors (a session with a `parentSessionId`) stay closeable. Both lists
465
+ are workflow inputs; a trailing `*` is a prefix wildcard.
466
+ - The report carries an `origin` column and the retained action (e.g.
467
+ `flag (origine utilisateur)`), in dry run as well as apply.
261
468
 
262
469
  ## The judge
263
470