@agentproto/apps 0.18.0 → 0.20.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/index.d.ts +2 -0
- package/dist/index.d.ts.map +1 -1
- package/dist/index.mjs +61 -2
- package/dist/index.mjs.map +1 -1
- package/dist/review-panel/panel.d.ts +1 -1
- package/dist/review-panel/panel.d.ts.map +1 -1
- package/dist/review-panel/panel.generated.d.ts +1 -1
- package/dist/review-panel/panel.generated.d.ts.map +1 -1
- package/dist/review-panel/panel.mjs +1 -1
- package/dist/review-panel/panel.mjs.map +1 -1
- package/dist/review-panel.mjs +1 -1
- package/dist/review-panel.mjs.map +1 -1
- package/dist/store/index.d.ts +61 -0
- package/dist/store/index.d.ts.map +1 -0
- package/dist/store/panel.d.ts +53 -0
- package/dist/store/panel.d.ts.map +1 -0
- package/dist/store/panel.generated.d.ts +13 -0
- package/dist/store/panel.generated.d.ts.map +1 -0
- package/dist/store/panel.mjs +25 -0
- package/dist/store/panel.mjs.map +1 -0
- package/dist/store.mjs +71 -0
- package/dist/store.mjs.map +1 -0
- package/package.json +14 -2
- package/session-steward/.agentproto/APP.md +2 -0
- package/session-steward/.agentproto/workflows/session-steward/WORKFLOW.md +218 -11
- package/session-steward/.agentproto/workflows/session-steward/cron-rules.mjs +526 -0
- package/session-steward/.agentproto/workflows/session-steward/entry.mjs +539 -64
- package/session-steward/.agentproto/workflows/session-steward/origin-policy.mjs +115 -0
- package/session-steward/README.md +19 -2
- package/session-steward/routines/session-steward-hourly/ROUTINE.md +24 -1
- package/session-steward/scripts/sessions-snapshot.sh +85 -0
- package/session-steward/skill/SKILL.md +74 -0
|
@@ -53,6 +53,43 @@ inputs:
|
|
|
53
53
|
callerSessionId:
|
|
54
54
|
type: string
|
|
55
55
|
description: The calling session's id — never a candidate.
|
|
56
|
+
callerOrigin:
|
|
57
|
+
type: string
|
|
58
|
+
description: >-
|
|
59
|
+
The calling session's origin (`cron:<jobId>`) — an older run of the
|
|
60
|
+
SAME cron job is never judged as user work.
|
|
61
|
+
appId:
|
|
62
|
+
type: string
|
|
63
|
+
description: >-
|
|
64
|
+
Installed app whose `app_state` ledger holds the verdict memory.
|
|
65
|
+
Default `@agentproto/session-steward` (the id in APP.md). A bare name
|
|
66
|
+
also matches a scoped install; a missing app means no memory, noted in
|
|
67
|
+
the report.
|
|
68
|
+
stableVerdictPasses:
|
|
69
|
+
type: number
|
|
70
|
+
description: >-
|
|
71
|
+
Consecutive passes on an unchanged evidence fingerprint before the
|
|
72
|
+
judge cache stops re-judging a session.
|
|
73
|
+
default: 2
|
|
74
|
+
userOrigins:
|
|
75
|
+
type: array
|
|
76
|
+
description: >-
|
|
77
|
+
Origins that are ALWAYS flag-only, never closed (a human is in the
|
|
78
|
+
loop). A trailing `*` is a prefix wildcard. Default
|
|
79
|
+
`["chat-starter", "vscode"]`; a root with no origin and no parent is
|
|
80
|
+
treated as a user origin too.
|
|
81
|
+
items:
|
|
82
|
+
type: string
|
|
83
|
+
default: ["chat-starter", "vscode"]
|
|
84
|
+
closableOrigins:
|
|
85
|
+
type: array
|
|
86
|
+
description: >-
|
|
87
|
+
Origins that may be closed under the current rules. A trailing `*` is a
|
|
88
|
+
prefix wildcard. Default `["cron:*", "gate"]`. Executors (a session with
|
|
89
|
+
a `parentSessionId`) are closable regardless.
|
|
90
|
+
items:
|
|
91
|
+
type: string
|
|
92
|
+
default: ["cron:*", "gate"]
|
|
56
93
|
outputs: {}
|
|
57
94
|
steps:
|
|
58
95
|
- id: modelRoles
|
|
@@ -75,6 +112,7 @@ steps:
|
|
|
75
112
|
tool: session_wrapup_plan
|
|
76
113
|
inputs:
|
|
77
114
|
idleMinutes: $steps.settings.idleMinutes
|
|
115
|
+
wait: true
|
|
78
116
|
|
|
79
117
|
- id: candidates
|
|
80
118
|
kind: transform
|
|
@@ -83,11 +121,36 @@ steps:
|
|
|
83
121
|
Entry-based — splitCandidates. `judge` is ordered most RAM first and
|
|
84
122
|
capped at `maxJudged`.
|
|
85
123
|
|
|
124
|
+
- id: hostLoad
|
|
125
|
+
kind: tool
|
|
126
|
+
name: Host saturation report (report only)
|
|
127
|
+
tool: host_load
|
|
128
|
+
inputs: {}
|
|
129
|
+
|
|
130
|
+
- id: liveSessions
|
|
131
|
+
kind: tool
|
|
132
|
+
name: List live sessions for the mechanical scan
|
|
133
|
+
tool: session_list
|
|
134
|
+
inputs:
|
|
135
|
+
full: true
|
|
136
|
+
|
|
137
|
+
- id: scan
|
|
138
|
+
kind: transform
|
|
139
|
+
name: Scan live sessions (busy / idle / terminal / never-ran / excluded)
|
|
140
|
+
description: Entry-based — scanLive.
|
|
141
|
+
|
|
142
|
+
- id: candidatesPlus
|
|
143
|
+
kind: transform
|
|
144
|
+
name: Merge never-ran 0/0 sessions into stuck
|
|
145
|
+
description: Entry-based — mergeNeverRan (never-ran is stuck, never judged).
|
|
146
|
+
|
|
86
147
|
- id: ruleApplyQueue
|
|
87
148
|
kind: transform
|
|
88
149
|
name: Rule verdicts to apply
|
|
89
150
|
description: >-
|
|
90
151
|
Entry-based. Empty unless `apply`: `close` → done, `stuck` → abandoned.
|
|
152
|
+
Origin-bounded — a user-origin candidate is queued as a `needs-input`
|
|
153
|
+
FLAG instead of a close.
|
|
91
154
|
|
|
92
155
|
- id: autoApply
|
|
93
156
|
kind: map
|
|
@@ -104,10 +167,81 @@ steps:
|
|
|
104
167
|
verdict: $item.verdict
|
|
105
168
|
note: $item.note
|
|
106
169
|
|
|
170
|
+
- id: installedApps
|
|
171
|
+
kind: tool
|
|
172
|
+
name: List installed apps (to resolve the memory app)
|
|
173
|
+
tool: app_list
|
|
174
|
+
inputs: {}
|
|
175
|
+
|
|
176
|
+
- id: memoryApp
|
|
177
|
+
kind: transform
|
|
178
|
+
name: Resolve the installed app holding the verdict memory
|
|
179
|
+
description: Entry-based — resolveMemoryApp. No such app ⇒ no memory + a report note.
|
|
180
|
+
|
|
181
|
+
- id: memoryQueue
|
|
182
|
+
kind: transform
|
|
183
|
+
name: Memory read queue
|
|
184
|
+
description: Entry-based. Empty when no installed memory app resolves.
|
|
185
|
+
|
|
186
|
+
- id: memoryRead
|
|
187
|
+
kind: map
|
|
188
|
+
name: Read the verdict memory ledger (best-effort)
|
|
189
|
+
over: $steps.memoryQueue
|
|
190
|
+
parallelism: 1
|
|
191
|
+
onError: collect
|
|
192
|
+
steps:
|
|
193
|
+
- id: memoryReadOne
|
|
194
|
+
kind: tool
|
|
195
|
+
tool: app_state_list
|
|
196
|
+
inputs:
|
|
197
|
+
appId: $item.appId
|
|
198
|
+
stage: session-steward
|
|
199
|
+
kinds: [note]
|
|
200
|
+
limit: 500
|
|
201
|
+
|
|
202
|
+
- id: memory
|
|
203
|
+
kind: transform
|
|
204
|
+
name: Fold verdict memory
|
|
205
|
+
description: Entry-based — foldMemory.
|
|
206
|
+
|
|
207
|
+
- id: loopScan
|
|
208
|
+
kind: map
|
|
209
|
+
name: Loop sanity per busy session
|
|
210
|
+
over: $steps.scan.loopQueue
|
|
211
|
+
parallelism: 4
|
|
212
|
+
onError: collect
|
|
213
|
+
steps:
|
|
214
|
+
- id: loopCallsOne
|
|
215
|
+
kind: tool
|
|
216
|
+
tool: tool_calls_list
|
|
217
|
+
inputs:
|
|
218
|
+
sessionId: $item.sessionId
|
|
219
|
+
lastN: 60
|
|
220
|
+
- id: loopFold
|
|
221
|
+
kind: transform
|
|
222
|
+
name: Loop verdict for one session
|
|
223
|
+
|
|
224
|
+
- id: loopResults
|
|
225
|
+
kind: transform
|
|
226
|
+
name: Loop verdicts
|
|
227
|
+
description: Entry-based.
|
|
228
|
+
|
|
229
|
+
- id: proposals
|
|
230
|
+
kind: transform
|
|
231
|
+
name: Nudge proposals (report only — never executed here)
|
|
232
|
+
description: >-
|
|
233
|
+
Entry-based — buildProposalsStep: loop → interrupt, stall → continue, at
|
|
234
|
+
most one per session per pass; a user-origin session is observed only.
|
|
235
|
+
|
|
236
|
+
- id: relabelQueue
|
|
237
|
+
kind: transform
|
|
238
|
+
name: Terminal sessions missing an outcome (relabel candidates)
|
|
239
|
+
description: Entry-based — buildRelabelQueue.
|
|
240
|
+
|
|
107
241
|
- id: evidence
|
|
108
242
|
kind: map
|
|
109
243
|
name: Collect compact evidence per judge candidate
|
|
110
|
-
over: $steps.
|
|
244
|
+
over: $steps.candidatesPlus.judge
|
|
111
245
|
parallelism: 4
|
|
112
246
|
onError: collect
|
|
113
247
|
steps:
|
|
@@ -119,8 +253,8 @@ steps:
|
|
|
119
253
|
|
|
120
254
|
- id: judgeQueue
|
|
121
255
|
kind: transform
|
|
122
|
-
name: Candidates with evidence
|
|
123
|
-
description: Entry-based.
|
|
256
|
+
name: Candidates with evidence minus the stable-verdict cache
|
|
257
|
+
description: Entry-based — buildJudgeQueueFiltered.
|
|
124
258
|
|
|
125
259
|
- id: jevQueue
|
|
126
260
|
kind: transform
|
|
@@ -207,7 +341,8 @@ steps:
|
|
|
207
341
|
name: Confident verdicts to apply
|
|
208
342
|
description: >-
|
|
209
343
|
Entry-based. Empty unless `apply`: done/abandoned/blocked/needs-input at
|
|
210
|
-
or above `minConfidence`.
|
|
344
|
+
or above `minConfidence`. Origin-bounded — a user-origin candidate is
|
|
345
|
+
downgraded to a `needs-input` FLAG, never a close.
|
|
211
346
|
|
|
212
347
|
- id: judgedApply
|
|
213
348
|
kind: map
|
|
@@ -225,6 +360,25 @@ steps:
|
|
|
225
360
|
judgedBy: $item.judgedBy
|
|
226
361
|
note: $item.note
|
|
227
362
|
|
|
363
|
+
- id: memoryWriteQueue
|
|
364
|
+
kind: transform
|
|
365
|
+
name: Verdict memory events to append
|
|
366
|
+
description: Entry-based — buildMemoryWriteQueue (a ledger write, never a session action).
|
|
367
|
+
|
|
368
|
+
- id: memoryWrite
|
|
369
|
+
kind: map
|
|
370
|
+
name: Append verdict memory to app_state (best-effort)
|
|
371
|
+
over: $steps.memoryWriteQueue
|
|
372
|
+
parallelism: 1
|
|
373
|
+
onError: collect
|
|
374
|
+
steps:
|
|
375
|
+
- id: memoryWriteOne
|
|
376
|
+
kind: tool
|
|
377
|
+
tool: app_state_append
|
|
378
|
+
inputs:
|
|
379
|
+
appId: $item.appId
|
|
380
|
+
event: $item.event
|
|
381
|
+
|
|
228
382
|
- id: report
|
|
229
383
|
kind: transform
|
|
230
384
|
name: Build the markdown report
|
|
@@ -233,23 +387,67 @@ steps:
|
|
|
233
387
|
result:
|
|
234
388
|
report: $steps.report
|
|
235
389
|
apply: $steps.settings.apply
|
|
236
|
-
candidates: $steps.
|
|
390
|
+
candidates: $steps.candidatesPlus
|
|
237
391
|
verdicts: $steps.finalVerdicts
|
|
238
392
|
autoApply: $steps.autoApply
|
|
239
393
|
judgedApply: $steps.judgedApply
|
|
394
|
+
proposals: $steps.proposals
|
|
395
|
+
relabel: $steps.relabelQueue
|
|
396
|
+
scan: $steps.scan
|
|
240
397
|
---
|
|
241
398
|
|
|
242
399
|
# Session Steward — `session-steward` workflow
|
|
243
400
|
|
|
244
|
-
`session_wrapup_plan` →
|
|
245
|
-
|
|
246
|
-
close
|
|
247
|
-
|
|
401
|
+
`session_wrapup_plan` → deterministic mechanical pass (host saturation, loop /
|
|
402
|
+
stall / never-ran / terminal-relabel scan, verdict memory) → rules pass over
|
|
403
|
+
`close`/`stuck` → compact evidence per `judge` session → one cheap judge turn
|
|
404
|
+
each (cached when the evidence is unchanged) → (opt-in) ask the session itself
|
|
405
|
+
→ close or flag confident verdicts through `session_wrapup_apply` → markdown
|
|
406
|
+
report with proposals, relabel candidates, and RAM freed / still held.
|
|
407
|
+
|
|
408
|
+
## Mechanical rules (ported from the `kill-idle-sessions` cron prototype)
|
|
409
|
+
|
|
410
|
+
Every rule below is a pure function in `cron-rules.mjs`, pinned by
|
|
411
|
+
`session-steward-cron-rules.test.ts`:
|
|
412
|
+
|
|
413
|
+
- **Loop (1).** `tool_calls_list` per busy session: the same argv verbatim ≥3
|
|
414
|
+
in 10 min, distinct/total < 0.2, or the same file read ≥4 → `looping`, a
|
|
415
|
+
sub-case of `active`. The proposed action is an **interrupt nudge**, never a
|
|
416
|
+
close. Useful loops (watch, test/type-check re-runs, `git status`, `gh pr`
|
|
417
|
+
polling) are excluded.
|
|
418
|
+
- **Stall (2).** Busy > 20 min with no new activity, or a recent
|
|
419
|
+
`lastTurnErroredAt` on an idle process → a proposed **"continue" nudge**.
|
|
420
|
+
- **Never-ran (3).** `tokensIn === 0 && tokensOut === 0` → `stuck`
|
|
421
|
+
immediately, without a judge, whatever the idle.
|
|
422
|
+
- **Fast-path done (4).** Last tool call is `message_parent(kind:done)` plus a
|
|
423
|
+
commit/PR → `done` without a judge (used by the criteria, see below).
|
|
424
|
+
- **Terminal without outcome (5).** Terminal sessions missing an outcome are
|
|
425
|
+
surfaced as relabel candidates instead of staying invisible.
|
|
426
|
+
- **Re-check at apply (6).** A candidate that became busy before the apply is
|
|
427
|
+
skipped (the apply tool also re-classifies).
|
|
428
|
+
- **Self-exclusion (7).** An older run of the caller's own `cron:<job>` is
|
|
429
|
+
never judged as user work.
|
|
430
|
+
- **Explicit 0-candidate report (8).** When nothing is idle, the report says
|
|
431
|
+
why (`n live, m busy, k terminal, j excluded`).
|
|
432
|
+
- **Host saturation (9).** If `host_load` is critical, the report lists
|
|
433
|
+
orphans and big non-session processes FIRST — report only, no action.
|
|
434
|
+
- **Verdict memory (10).** Each verdict is written to the app's `app_state`
|
|
435
|
+
ledger; a session judged the same verdict on an unchanged evidence
|
|
436
|
+
fingerprint for `stableVerdictPasses` passes is served from cache and not
|
|
437
|
+
re-judged. Operator disagreements are recorded as examples.
|
|
438
|
+
|
|
439
|
+
The proposed nudges (interrupt / continue) are **reported only** — this
|
|
440
|
+
workflow never sends a prompt and never closes a `looping` session. At most
|
|
441
|
+
one nudge is proposed per session per pass; a user-origin session is reported
|
|
442
|
+
as observed, never nudged.
|
|
248
443
|
|
|
249
444
|
## Safety
|
|
250
445
|
|
|
251
|
-
- `apply: false` (the default) mutates
|
|
252
|
-
an empty list.
|
|
446
|
+
- `apply: false` (the default) mutates no SESSION: every session-mutating map
|
|
447
|
+
runs over an empty list. The one write a dry run performs is the append-only
|
|
448
|
+
verdict-memory ledger (`app_state_append`) — it never touches a session and
|
|
449
|
+
is what lets the cache accumulate across passes. Set `appId: ""` to disable
|
|
450
|
+
it entirely.
|
|
253
451
|
- `session_wrapup_apply` re-classifies each id right before acting and always
|
|
254
452
|
refuses `keep`-class ids; this workflow never feeds it one.
|
|
255
453
|
- Rules only ever close `close`/`stuck` ids; a `keepAlive` session is never
|
|
@@ -258,6 +456,15 @@ report with RAM freed / still held.
|
|
|
258
456
|
- A malformed judge reply is `active` with confidence 0 — never acted on.
|
|
259
457
|
- The caller's own session (`callerSessionId`) is dropped from every list.
|
|
260
458
|
- `blocked` / `needs-input` only FLAG a session; it keeps running.
|
|
459
|
+
- **Origin bound (never close a human's session).** Every candidate's
|
|
460
|
+
`origin`/`parentSessionId` runs through the pure `decideAction`
|
|
461
|
+
(`origin-policy.mjs`): a `userOrigins` match (`chat-starter`, `vscode` by
|
|
462
|
+
default) or a root with no origin and no parent is FLAG-ONLY, even with
|
|
463
|
+
`apply: true` and a confident `done` verdict. `cron:*`, `gate`, and
|
|
464
|
+
executors (a session with a `parentSessionId`) stay closeable. Both lists
|
|
465
|
+
are workflow inputs; a trailing `*` is a prefix wildcard.
|
|
466
|
+
- The report carries an `origin` column and the retained action (e.g.
|
|
467
|
+
`flag (origine utilisateur)`), in dry run as well as apply.
|
|
261
468
|
|
|
262
469
|
## The judge
|
|
263
470
|
|