@agentproto/apps 0.18.0 → 0.19.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/index.mjs +1 -1
- package/dist/index.mjs.map +1 -1
- package/dist/review-panel/panel.d.ts +1 -1
- package/dist/review-panel/panel.d.ts.map +1 -1
- package/dist/review-panel/panel.generated.d.ts +1 -1
- package/dist/review-panel/panel.generated.d.ts.map +1 -1
- package/dist/review-panel/panel.mjs +1 -1
- package/dist/review-panel/panel.mjs.map +1 -1
- package/dist/review-panel.mjs +1 -1
- package/dist/review-panel.mjs.map +1 -1
- package/package.json +2 -2
- package/session-steward/.agentproto/APP.md +2 -0
- package/session-steward/.agentproto/workflows/session-steward/WORKFLOW.md +204 -11
- package/session-steward/.agentproto/workflows/session-steward/cron-rules.mjs +526 -0
- package/session-steward/.agentproto/workflows/session-steward/entry.mjs +510 -63
- package/session-steward/.agentproto/workflows/session-steward/origin-policy.mjs +115 -0
- package/session-steward/README.md +19 -2
- package/session-steward/routines/session-steward-hourly/ROUTINE.md +24 -1
- package/session-steward/scripts/sessions-snapshot.sh +85 -0
- package/session-steward/skill/SKILL.md +74 -0
|
@@ -53,6 +53,41 @@ inputs:
|
|
|
53
53
|
callerSessionId:
|
|
54
54
|
type: string
|
|
55
55
|
description: The calling session's id — never a candidate.
|
|
56
|
+
callerOrigin:
|
|
57
|
+
type: string
|
|
58
|
+
description: >-
|
|
59
|
+
The calling session's origin (`cron:<jobId>`) — an older run of the
|
|
60
|
+
SAME cron job is never judged as user work.
|
|
61
|
+
appId:
|
|
62
|
+
type: string
|
|
63
|
+
description: >-
|
|
64
|
+
Installed app whose `app_state` ledger holds the verdict memory.
|
|
65
|
+
Default `session-steward`.
|
|
66
|
+
stableVerdictPasses:
|
|
67
|
+
type: number
|
|
68
|
+
description: >-
|
|
69
|
+
Consecutive passes on an unchanged evidence fingerprint before the
|
|
70
|
+
judge cache stops re-judging a session.
|
|
71
|
+
default: 2
|
|
72
|
+
userOrigins:
|
|
73
|
+
type: array
|
|
74
|
+
description: >-
|
|
75
|
+
Origins that are ALWAYS flag-only, never closed (a human is in the
|
|
76
|
+
loop). A trailing `*` is a prefix wildcard. Default
|
|
77
|
+
`["chat-starter", "vscode"]`; a root with no origin and no parent is
|
|
78
|
+
treated as a user origin too.
|
|
79
|
+
items:
|
|
80
|
+
type: string
|
|
81
|
+
default: ["chat-starter", "vscode"]
|
|
82
|
+
closableOrigins:
|
|
83
|
+
type: array
|
|
84
|
+
description: >-
|
|
85
|
+
Origins that may be closed under the current rules. A trailing `*` is a
|
|
86
|
+
prefix wildcard. Default `["cron:*", "gate"]`. Executors (a session with
|
|
87
|
+
a `parentSessionId`) are closable regardless.
|
|
88
|
+
items:
|
|
89
|
+
type: string
|
|
90
|
+
default: ["cron:*", "gate"]
|
|
56
91
|
outputs: {}
|
|
57
92
|
steps:
|
|
58
93
|
- id: modelRoles
|
|
@@ -83,11 +118,36 @@ steps:
|
|
|
83
118
|
Entry-based — splitCandidates. `judge` is ordered most RAM first and
|
|
84
119
|
capped at `maxJudged`.
|
|
85
120
|
|
|
121
|
+
- id: hostLoad
|
|
122
|
+
kind: tool
|
|
123
|
+
name: Host saturation report (report only)
|
|
124
|
+
tool: host_load
|
|
125
|
+
inputs: {}
|
|
126
|
+
|
|
127
|
+
- id: liveSessions
|
|
128
|
+
kind: tool
|
|
129
|
+
name: List live sessions for the mechanical scan
|
|
130
|
+
tool: session_list
|
|
131
|
+
inputs:
|
|
132
|
+
full: true
|
|
133
|
+
|
|
134
|
+
- id: scan
|
|
135
|
+
kind: transform
|
|
136
|
+
name: Scan live sessions (busy / idle / terminal / never-ran / excluded)
|
|
137
|
+
description: Entry-based — scanLive.
|
|
138
|
+
|
|
139
|
+
- id: candidatesPlus
|
|
140
|
+
kind: transform
|
|
141
|
+
name: Merge never-ran 0/0 sessions into stuck
|
|
142
|
+
description: Entry-based — mergeNeverRan (never-ran is stuck, never judged).
|
|
143
|
+
|
|
86
144
|
- id: ruleApplyQueue
|
|
87
145
|
kind: transform
|
|
88
146
|
name: Rule verdicts to apply
|
|
89
147
|
description: >-
|
|
90
148
|
Entry-based. Empty unless `apply`: `close` → done, `stuck` → abandoned.
|
|
149
|
+
Origin-bounded — a user-origin candidate is queued as a `needs-input`
|
|
150
|
+
FLAG instead of a close.
|
|
91
151
|
|
|
92
152
|
- id: autoApply
|
|
93
153
|
kind: map
|
|
@@ -104,10 +164,70 @@ steps:
|
|
|
104
164
|
verdict: $item.verdict
|
|
105
165
|
note: $item.note
|
|
106
166
|
|
|
167
|
+
- id: memoryQueue
|
|
168
|
+
kind: transform
|
|
169
|
+
name: Memory read queue
|
|
170
|
+
description: Entry-based. Empty when no `appId` is configured.
|
|
171
|
+
|
|
172
|
+
- id: memoryRead
|
|
173
|
+
kind: map
|
|
174
|
+
name: Read the verdict memory ledger (best-effort)
|
|
175
|
+
over: $steps.memoryQueue
|
|
176
|
+
parallelism: 1
|
|
177
|
+
onError: collect
|
|
178
|
+
steps:
|
|
179
|
+
- id: memoryReadOne
|
|
180
|
+
kind: tool
|
|
181
|
+
tool: app_state_list
|
|
182
|
+
inputs:
|
|
183
|
+
appId: $item.appId
|
|
184
|
+
stage: session-steward
|
|
185
|
+
kinds: [note]
|
|
186
|
+
limit: 500
|
|
187
|
+
|
|
188
|
+
- id: memory
|
|
189
|
+
kind: transform
|
|
190
|
+
name: Fold verdict memory
|
|
191
|
+
description: Entry-based — foldMemory.
|
|
192
|
+
|
|
193
|
+
- id: loopScan
|
|
194
|
+
kind: map
|
|
195
|
+
name: Loop sanity per busy session
|
|
196
|
+
over: $steps.scan.loopQueue
|
|
197
|
+
parallelism: 4
|
|
198
|
+
onError: collect
|
|
199
|
+
steps:
|
|
200
|
+
- id: loopCallsOne
|
|
201
|
+
kind: tool
|
|
202
|
+
tool: tool_calls_list
|
|
203
|
+
inputs:
|
|
204
|
+
sessionId: $item.sessionId
|
|
205
|
+
lastN: 60
|
|
206
|
+
- id: loopFold
|
|
207
|
+
kind: transform
|
|
208
|
+
name: Loop verdict for one session
|
|
209
|
+
|
|
210
|
+
- id: loopResults
|
|
211
|
+
kind: transform
|
|
212
|
+
name: Loop verdicts
|
|
213
|
+
description: Entry-based.
|
|
214
|
+
|
|
215
|
+
- id: proposals
|
|
216
|
+
kind: transform
|
|
217
|
+
name: Nudge proposals (report only — never executed here)
|
|
218
|
+
description: >-
|
|
219
|
+
Entry-based — buildProposalsStep: loop → interrupt, stall → continue, at
|
|
220
|
+
most one per session per pass; a user-origin session is observed only.
|
|
221
|
+
|
|
222
|
+
- id: relabelQueue
|
|
223
|
+
kind: transform
|
|
224
|
+
name: Terminal sessions missing an outcome (relabel candidates)
|
|
225
|
+
description: Entry-based — buildRelabelQueue.
|
|
226
|
+
|
|
107
227
|
- id: evidence
|
|
108
228
|
kind: map
|
|
109
229
|
name: Collect compact evidence per judge candidate
|
|
110
|
-
over: $steps.
|
|
230
|
+
over: $steps.candidatesPlus.judge
|
|
111
231
|
parallelism: 4
|
|
112
232
|
onError: collect
|
|
113
233
|
steps:
|
|
@@ -119,8 +239,8 @@ steps:
|
|
|
119
239
|
|
|
120
240
|
- id: judgeQueue
|
|
121
241
|
kind: transform
|
|
122
|
-
name: Candidates with evidence
|
|
123
|
-
description: Entry-based.
|
|
242
|
+
name: Candidates with evidence minus the stable-verdict cache
|
|
243
|
+
description: Entry-based — buildJudgeQueueFiltered.
|
|
124
244
|
|
|
125
245
|
- id: jevQueue
|
|
126
246
|
kind: transform
|
|
@@ -207,7 +327,8 @@ steps:
|
|
|
207
327
|
name: Confident verdicts to apply
|
|
208
328
|
description: >-
|
|
209
329
|
Entry-based. Empty unless `apply`: done/abandoned/blocked/needs-input at
|
|
210
|
-
or above `minConfidence`.
|
|
330
|
+
or above `minConfidence`. Origin-bounded — a user-origin candidate is
|
|
331
|
+
downgraded to a `needs-input` FLAG, never a close.
|
|
211
332
|
|
|
212
333
|
- id: judgedApply
|
|
213
334
|
kind: map
|
|
@@ -225,6 +346,25 @@ steps:
|
|
|
225
346
|
judgedBy: $item.judgedBy
|
|
226
347
|
note: $item.note
|
|
227
348
|
|
|
349
|
+
- id: memoryWriteQueue
|
|
350
|
+
kind: transform
|
|
351
|
+
name: Verdict memory events to append
|
|
352
|
+
description: Entry-based — buildMemoryWriteQueue (a ledger write, never a session action).
|
|
353
|
+
|
|
354
|
+
- id: memoryWrite
|
|
355
|
+
kind: map
|
|
356
|
+
name: Append verdict memory to app_state (best-effort)
|
|
357
|
+
over: $steps.memoryWriteQueue
|
|
358
|
+
parallelism: 1
|
|
359
|
+
onError: collect
|
|
360
|
+
steps:
|
|
361
|
+
- id: memoryWriteOne
|
|
362
|
+
kind: tool
|
|
363
|
+
tool: app_state_append
|
|
364
|
+
inputs:
|
|
365
|
+
appId: $item.appId
|
|
366
|
+
event: $item.event
|
|
367
|
+
|
|
228
368
|
- id: report
|
|
229
369
|
kind: transform
|
|
230
370
|
name: Build the markdown report
|
|
@@ -233,23 +373,67 @@ steps:
|
|
|
233
373
|
result:
|
|
234
374
|
report: $steps.report
|
|
235
375
|
apply: $steps.settings.apply
|
|
236
|
-
candidates: $steps.
|
|
376
|
+
candidates: $steps.candidatesPlus
|
|
237
377
|
verdicts: $steps.finalVerdicts
|
|
238
378
|
autoApply: $steps.autoApply
|
|
239
379
|
judgedApply: $steps.judgedApply
|
|
380
|
+
proposals: $steps.proposals
|
|
381
|
+
relabel: $steps.relabelQueue
|
|
382
|
+
scan: $steps.scan
|
|
240
383
|
---
|
|
241
384
|
|
|
242
385
|
# Session Steward — `session-steward` workflow
|
|
243
386
|
|
|
244
|
-
`session_wrapup_plan` →
|
|
245
|
-
|
|
246
|
-
close
|
|
247
|
-
|
|
387
|
+
`session_wrapup_plan` → deterministic mechanical pass (host saturation, loop /
|
|
388
|
+
stall / never-ran / terminal-relabel scan, verdict memory) → rules pass over
|
|
389
|
+
`close`/`stuck` → compact evidence per `judge` session → one cheap judge turn
|
|
390
|
+
each (cached when the evidence is unchanged) → (opt-in) ask the session itself
|
|
391
|
+
→ close or flag confident verdicts through `session_wrapup_apply` → markdown
|
|
392
|
+
report with proposals, relabel candidates, and RAM freed / still held.
|
|
393
|
+
|
|
394
|
+
## Mechanical rules (ported from the `kill-idle-sessions` cron prototype)
|
|
395
|
+
|
|
396
|
+
Every rule below is a pure function in `cron-rules.mjs`, pinned by
|
|
397
|
+
`session-steward-cron-rules.test.ts`:
|
|
398
|
+
|
|
399
|
+
- **Loop (1).** `tool_calls_list` per busy session: the same argv verbatim ≥3
|
|
400
|
+
in 10 min, distinct/total < 0.2, or the same file read ≥4 → `looping`, a
|
|
401
|
+
sub-case of `active`. The proposed action is an **interrupt nudge**, never a
|
|
402
|
+
close. Useful loops (watch, test/type-check re-runs, `git status`, `gh pr`
|
|
403
|
+
polling) are excluded.
|
|
404
|
+
- **Stall (2).** Busy > 20 min with no new activity, or a recent
|
|
405
|
+
`lastTurnErroredAt` on an idle process → a proposed **"continue" nudge**.
|
|
406
|
+
- **Never-ran (3).** `tokensIn === 0 && tokensOut === 0` → `stuck`
|
|
407
|
+
immediately, without a judge, whatever the idle.
|
|
408
|
+
- **Fast-path done (4).** Last tool call is `message_parent(kind:done)` plus a
|
|
409
|
+
commit/PR → `done` without a judge (used by the criteria, see below).
|
|
410
|
+
- **Terminal without outcome (5).** Terminal sessions missing an outcome are
|
|
411
|
+
surfaced as relabel candidates instead of staying invisible.
|
|
412
|
+
- **Re-check at apply (6).** A candidate that became busy before the apply is
|
|
413
|
+
skipped (the apply tool also re-classifies).
|
|
414
|
+
- **Self-exclusion (7).** An older run of the caller's own `cron:<job>` is
|
|
415
|
+
never judged as user work.
|
|
416
|
+
- **Explicit 0-candidate report (8).** When nothing is idle, the report says
|
|
417
|
+
why (`n live, m busy, k terminal, j excluded`).
|
|
418
|
+
- **Host saturation (9).** If `host_load` is critical, the report lists
|
|
419
|
+
orphans and big non-session processes FIRST — report only, no action.
|
|
420
|
+
- **Verdict memory (10).** Each verdict is written to the app's `app_state`
|
|
421
|
+
ledger; a session judged the same verdict on an unchanged evidence
|
|
422
|
+
fingerprint for `stableVerdictPasses` passes is served from cache and not
|
|
423
|
+
re-judged. Operator disagreements are recorded as examples.
|
|
424
|
+
|
|
425
|
+
The proposed nudges (interrupt / continue) are **reported only** — this
|
|
426
|
+
workflow never sends a prompt and never closes a `looping` session. At most
|
|
427
|
+
one nudge is proposed per session per pass; a user-origin session is reported
|
|
428
|
+
as observed, never nudged.
|
|
248
429
|
|
|
249
430
|
## Safety
|
|
250
431
|
|
|
251
|
-
- `apply: false` (the default) mutates
|
|
252
|
-
an empty list.
|
|
432
|
+
- `apply: false` (the default) mutates no SESSION: every session-mutating map
|
|
433
|
+
runs over an empty list. The one write a dry run performs is the append-only
|
|
434
|
+
verdict-memory ledger (`app_state_append`) — it never touches a session and
|
|
435
|
+
is what lets the cache accumulate across passes. Set `appId: ""` to disable
|
|
436
|
+
it entirely.
|
|
253
437
|
- `session_wrapup_apply` re-classifies each id right before acting and always
|
|
254
438
|
refuses `keep`-class ids; this workflow never feeds it one.
|
|
255
439
|
- Rules only ever close `close`/`stuck` ids; a `keepAlive` session is never
|
|
@@ -258,6 +442,15 @@ report with RAM freed / still held.
|
|
|
258
442
|
- A malformed judge reply is `active` with confidence 0 — never acted on.
|
|
259
443
|
- The caller's own session (`callerSessionId`) is dropped from every list.
|
|
260
444
|
- `blocked` / `needs-input` only FLAG a session; it keeps running.
|
|
445
|
+
- **Origin bound (never close a human's session).** Every candidate's
|
|
446
|
+
`origin`/`parentSessionId` runs through the pure `decideAction`
|
|
447
|
+
(`origin-policy.mjs`): a `userOrigins` match (`chat-starter`, `vscode` by
|
|
448
|
+
default) or a root with no origin and no parent is FLAG-ONLY, even with
|
|
449
|
+
`apply: true` and a confident `done` verdict. `cron:*`, `gate`, and
|
|
450
|
+
executors (a session with a `parentSessionId`) stay closeable. Both lists
|
|
451
|
+
are workflow inputs; a trailing `*` is a prefix wildcard.
|
|
452
|
+
- The report carries an `origin` column and the retained action (e.g.
|
|
453
|
+
`flag (origine utilisateur)`), in dry run as well as apply.
|
|
261
454
|
|
|
262
455
|
## The judge
|
|
263
456
|
|