@agentproto/apps 0.17.2 → 0.19.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -53,6 +53,41 @@ inputs:
53
53
  callerSessionId:
54
54
  type: string
55
55
  description: The calling session's id — never a candidate.
56
+ callerOrigin:
57
+ type: string
58
+ description: >-
59
+ The calling session's origin (`cron:<jobId>`) — an older run of the
60
+ SAME cron job is never judged as user work.
61
+ appId:
62
+ type: string
63
+ description: >-
64
+ Installed app whose `app_state` ledger holds the verdict memory.
65
+ Default `session-steward`.
66
+ stableVerdictPasses:
67
+ type: number
68
+ description: >-
69
+ Consecutive passes on an unchanged evidence fingerprint before the
70
+ judge cache stops re-judging a session.
71
+ default: 2
72
+ userOrigins:
73
+ type: array
74
+ description: >-
75
+ Origins that are ALWAYS flag-only, never closed (a human is in the
76
+ loop). A trailing `*` is a prefix wildcard. Default
77
+ `["chat-starter", "vscode"]`; a root with no origin and no parent is
78
+ treated as a user origin too.
79
+ items:
80
+ type: string
81
+ default: ["chat-starter", "vscode"]
82
+ closableOrigins:
83
+ type: array
84
+ description: >-
85
+ Origins that may be closed under the current rules. A trailing `*` is a
86
+ prefix wildcard. Default `["cron:*", "gate"]`. Executors (a session with
87
+ a `parentSessionId`) are closable regardless.
88
+ items:
89
+ type: string
90
+ default: ["cron:*", "gate"]
56
91
  outputs: {}
57
92
  steps:
58
93
  - id: modelRoles
@@ -83,11 +118,36 @@ steps:
83
118
  Entry-based — splitCandidates. `judge` is ordered most RAM first and
84
119
  capped at `maxJudged`.
85
120
 
121
+ - id: hostLoad
122
+ kind: tool
123
+ name: Host saturation report (report only)
124
+ tool: host_load
125
+ inputs: {}
126
+
127
+ - id: liveSessions
128
+ kind: tool
129
+ name: List live sessions for the mechanical scan
130
+ tool: session_list
131
+ inputs:
132
+ full: true
133
+
134
+ - id: scan
135
+ kind: transform
136
+ name: Scan live sessions (busy / idle / terminal / never-ran / excluded)
137
+ description: Entry-based — scanLive.
138
+
139
+ - id: candidatesPlus
140
+ kind: transform
141
+ name: Merge never-ran 0/0 sessions into stuck
142
+ description: Entry-based — mergeNeverRan (never-ran is stuck, never judged).
143
+
86
144
  - id: ruleApplyQueue
87
145
  kind: transform
88
146
  name: Rule verdicts to apply
89
147
  description: >-
90
148
  Entry-based. Empty unless `apply`: `close` → done, `stuck` → abandoned.
149
+ Origin-bounded — a user-origin candidate is queued as a `needs-input`
150
+ FLAG instead of a close.
91
151
 
92
152
  - id: autoApply
93
153
  kind: map
@@ -104,10 +164,70 @@ steps:
104
164
  verdict: $item.verdict
105
165
  note: $item.note
106
166
 
167
+ - id: memoryQueue
168
+ kind: transform
169
+ name: Memory read queue
170
+ description: Entry-based. Empty when no `appId` is configured.
171
+
172
+ - id: memoryRead
173
+ kind: map
174
+ name: Read the verdict memory ledger (best-effort)
175
+ over: $steps.memoryQueue
176
+ parallelism: 1
177
+ onError: collect
178
+ steps:
179
+ - id: memoryReadOne
180
+ kind: tool
181
+ tool: app_state_list
182
+ inputs:
183
+ appId: $item.appId
184
+ stage: session-steward
185
+ kinds: [note]
186
+ limit: 500
187
+
188
+ - id: memory
189
+ kind: transform
190
+ name: Fold verdict memory
191
+ description: Entry-based — foldMemory.
192
+
193
+ - id: loopScan
194
+ kind: map
195
+ name: Loop sanity per busy session
196
+ over: $steps.scan.loopQueue
197
+ parallelism: 4
198
+ onError: collect
199
+ steps:
200
+ - id: loopCallsOne
201
+ kind: tool
202
+ tool: tool_calls_list
203
+ inputs:
204
+ sessionId: $item.sessionId
205
+ lastN: 60
206
+ - id: loopFold
207
+ kind: transform
208
+ name: Loop verdict for one session
209
+
210
+ - id: loopResults
211
+ kind: transform
212
+ name: Loop verdicts
213
+ description: Entry-based.
214
+
215
+ - id: proposals
216
+ kind: transform
217
+ name: Nudge proposals (report only — never executed here)
218
+ description: >-
219
+ Entry-based — buildProposalsStep: loop → interrupt, stall → continue, at
220
+ most one per session per pass; a user-origin session is observed only.
221
+
222
+ - id: relabelQueue
223
+ kind: transform
224
+ name: Terminal sessions missing an outcome (relabel candidates)
225
+ description: Entry-based — buildRelabelQueue.
226
+
107
227
  - id: evidence
108
228
  kind: map
109
229
  name: Collect compact evidence per judge candidate
110
- over: $steps.candidates.judge
230
+ over: $steps.candidatesPlus.judge
111
231
  parallelism: 4
112
232
  onError: collect
113
233
  steps:
@@ -119,8 +239,8 @@ steps:
119
239
 
120
240
  - id: judgeQueue
121
241
  kind: transform
122
- name: Candidates with evidence
123
- description: Entry-based.
242
+ name: Candidates with evidence minus the stable-verdict cache
243
+ description: Entry-based — buildJudgeQueueFiltered.
124
244
 
125
245
  - id: jevQueue
126
246
  kind: transform
@@ -207,7 +327,8 @@ steps:
207
327
  name: Confident verdicts to apply
208
328
  description: >-
209
329
  Entry-based. Empty unless `apply`: done/abandoned/blocked/needs-input at
210
- or above `minConfidence`.
330
+ or above `minConfidence`. Origin-bounded — a user-origin candidate is
331
+ downgraded to a `needs-input` FLAG, never a close.
211
332
 
212
333
  - id: judgedApply
213
334
  kind: map
@@ -225,6 +346,25 @@ steps:
225
346
  judgedBy: $item.judgedBy
226
347
  note: $item.note
227
348
 
349
+ - id: memoryWriteQueue
350
+ kind: transform
351
+ name: Verdict memory events to append
352
+ description: Entry-based — buildMemoryWriteQueue (a ledger write, never a session action).
353
+
354
+ - id: memoryWrite
355
+ kind: map
356
+ name: Append verdict memory to app_state (best-effort)
357
+ over: $steps.memoryWriteQueue
358
+ parallelism: 1
359
+ onError: collect
360
+ steps:
361
+ - id: memoryWriteOne
362
+ kind: tool
363
+ tool: app_state_append
364
+ inputs:
365
+ appId: $item.appId
366
+ event: $item.event
367
+
228
368
  - id: report
229
369
  kind: transform
230
370
  name: Build the markdown report
@@ -233,23 +373,67 @@ steps:
233
373
  result:
234
374
  report: $steps.report
235
375
  apply: $steps.settings.apply
236
- candidates: $steps.candidates
376
+ candidates: $steps.candidatesPlus
237
377
  verdicts: $steps.finalVerdicts
238
378
  autoApply: $steps.autoApply
239
379
  judgedApply: $steps.judgedApply
380
+ proposals: $steps.proposals
381
+ relabel: $steps.relabelQueue
382
+ scan: $steps.scan
240
383
  ---
241
384
 
242
385
  # Session Steward — `session-steward` workflow
243
386
 
244
- `session_wrapup_plan` → rules pass over `close`/`stuck` → compact evidence per
245
- `judge` session → one cheap judge turn each → (opt-in) ask the session itself →
246
- close or flag confident verdicts through `session_wrapup_apply` → markdown
247
- report with RAM freed / still held.
387
+ `session_wrapup_plan` → deterministic mechanical pass (host saturation, loop /
388
+ stall / never-ran / terminal-relabel scan, verdict memory) → rules pass over
389
+ `close`/`stuck` → compact evidence per `judge` session → one cheap judge turn
390
+ each (cached when the evidence is unchanged) → (opt-in) ask the session itself
391
+ → close or flag confident verdicts through `session_wrapup_apply` → markdown
392
+ report with proposals, relabel candidates, and RAM freed / still held.
393
+
394
+ ## Mechanical rules (ported from the `kill-idle-sessions` cron prototype)
395
+
396
+ Every rule below is a pure function in `cron-rules.mjs`, pinned by
397
+ `session-steward-cron-rules.test.ts`:
398
+
399
+ - **Loop (1).** `tool_calls_list` per busy session: the same argv verbatim ≥3
400
+ in 10 min, distinct/total < 0.2, or the same file read ≥4 → `looping`, a
401
+ sub-case of `active`. The proposed action is an **interrupt nudge**, never a
402
+ close. Useful loops (watch, test/type-check re-runs, `git status`, `gh pr`
403
+ polling) are excluded.
404
+ - **Stall (2).** Busy > 20 min with no new activity, or a recent
405
+ `lastTurnErroredAt` on an idle process → a proposed **"continue" nudge**.
406
+ - **Never-ran (3).** `tokensIn === 0 && tokensOut === 0` → `stuck`
407
+ immediately, without a judge, whatever the idle.
408
+ - **Fast-path done (4).** Last tool call is `message_parent(kind:done)` plus a
409
+ commit/PR → `done` without a judge (used by the criteria, see below).
410
+ - **Terminal without outcome (5).** Terminal sessions missing an outcome are
411
+ surfaced as relabel candidates instead of staying invisible.
412
+ - **Re-check at apply (6).** A candidate that became busy before the apply is
413
+ skipped (the apply tool also re-classifies).
414
+ - **Self-exclusion (7).** An older run of the caller's own `cron:<job>` is
415
+ never judged as user work.
416
+ - **Explicit 0-candidate report (8).** When nothing is idle, the report says
417
+ why (`n live, m busy, k terminal, j excluded`).
418
+ - **Host saturation (9).** If `host_load` is critical, the report lists
419
+ orphans and big non-session processes FIRST — report only, no action.
420
+ - **Verdict memory (10).** Each verdict is written to the app's `app_state`
421
+ ledger; a session judged the same verdict on an unchanged evidence
422
+ fingerprint for `stableVerdictPasses` passes is served from cache and not
423
+ re-judged. Operator disagreements are recorded as examples.
424
+
425
+ The proposed nudges (interrupt / continue) are **reported only** — this
426
+ workflow never sends a prompt and never closes a `looping` session. At most
427
+ one nudge is proposed per session per pass; a user-origin session is reported
428
+ as observed, never nudged.
248
429
 
249
430
  ## Safety
250
431
 
251
- - `apply: false` (the default) mutates nothing: every mutating map runs over
252
- an empty list.
432
+ - `apply: false` (the default) mutates no SESSION: every session-mutating map
433
+ runs over an empty list. The one write a dry run performs is the append-only
434
+ verdict-memory ledger (`app_state_append`) — it never touches a session and
435
+ is what lets the cache accumulate across passes. Set `appId: ""` to disable
436
+ it entirely.
253
437
  - `session_wrapup_apply` re-classifies each id right before acting and always
254
438
  refuses `keep`-class ids; this workflow never feeds it one.
255
439
  - Rules only ever close `close`/`stuck` ids; a `keepAlive` session is never
@@ -258,6 +442,15 @@ report with RAM freed / still held.
258
442
  - A malformed judge reply is `active` with confidence 0 — never acted on.
259
443
  - The caller's own session (`callerSessionId`) is dropped from every list.
260
444
  - `blocked` / `needs-input` only FLAG a session; it keeps running.
445
+ - **Origin bound (never close a human's session).** Every candidate's
446
+ `origin`/`parentSessionId` runs through the pure `decideAction`
447
+ (`origin-policy.mjs`): a `userOrigins` match (`chat-starter`, `vscode` by
448
+ default) or a root with no origin and no parent is FLAG-ONLY, even with
449
+ `apply: true` and a confident `done` verdict. `cron:*`, `gate`, and
450
+ executors (a session with a `parentSessionId`) stay closeable. Both lists
451
+ are workflow inputs; a trailing `*` is a prefix wildcard.
452
+ - The report carries an `origin` column and the retained action (e.g.
453
+ `flag (origine utilisateur)`), in dry run as well as apply.
261
454
 
262
455
  ## The judge
263
456