pi-goal-list-loop-audit 0.35.4 → 0.35.13
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +7666 -0
- package/INSTALL.md +377 -0
- package/README.md +26 -13
- package/docs/DESIGN.md +31 -4
- package/docs/INDEX.md +32 -7
- package/examples/example-objective.md +119 -0
- package/extensions/faulty-objective-recovery.ts +13 -2
- package/extensions/goal-commands.ts +50 -6
- package/extensions/goal-heartbeat.ts +25 -21
- package/extensions/goal-loop-auditor-process.ts +184 -3
- package/extensions/goal-loop-core.ts +108 -15
- package/extensions/goal-loop-display.ts +82 -0
- package/extensions/goal-loop-shield.ts +55 -0
- package/extensions/goal-recovery.ts +317 -7
- package/extensions/goal-settings.ts +31 -1
- package/extensions/loops/goal-activation.ts +74 -7
- package/extensions/loops/goal-auditor-hooks.ts +30 -55
- package/extensions/loops/goal-list-queue.ts +18 -5
- package/extensions/loops/goal-orchestrator.ts +5 -0
- package/extensions/loops/goal-runtime-globals.ts +2 -0
- package/extensions/loops/goal-session.ts +101 -21
- package/extensions/loops/goal-settings-ui.ts +122 -63
- package/extensions/loops/goal-tools.ts +110 -36
- package/extensions/loops/goal-ui.ts +59 -2
- package/extensions/main-model-recovery.ts +24 -1
- package/extensions/model-picker.ts +3 -1
- package/extensions/model-selector.ts +2 -1
- package/extensions/multi-model-picker.ts +64 -7
- package/extensions/settings-menu.ts +16 -0
- package/package.json +4 -1
- package/prompts/goal-loop-continuation.md +11 -8
- package/prompts/goal-loop-draft.md +5 -5
- package/schemas/goal.schema.json +17 -0
package/INSTALL.md
ADDED
|
@@ -0,0 +1,377 @@
|
|
|
1
|
+
# Install & try
|
|
2
|
+
|
|
3
|
+
## Install (the usual way)
|
|
4
|
+
|
|
5
|
+
```bash
|
|
6
|
+
pi install npm:pi-goal-list-loop-audit
|
|
7
|
+
```
|
|
8
|
+
|
|
9
|
+
That's it — pi loads the extension into every session (run `/reload` in any
|
|
10
|
+
session that was already open).
|
|
11
|
+
|
|
12
|
+
> **Persistence note**: `pi update` can overwrite `~/.pi/agent/npm/node_modules/`.
|
|
13
|
+
> If the plugin disappears after an update, re-run `pi install`. For a permanent
|
|
14
|
+
> install, copy the package into your project's `.pi/extensions/` directory instead.
|
|
15
|
+
|
|
16
|
+
## Your first goal in 60 seconds
|
|
17
|
+
|
|
18
|
+
In any pi session, in the project you want worked on:
|
|
19
|
+
|
|
20
|
+
```
|
|
21
|
+
/goal "Fix the login timeout bug. Done when: `npm test` passes"
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
What happens:
|
|
25
|
+
|
|
26
|
+
1. **Draft + Confirm** — glla shows the goal contract (objective + Done-when);
|
|
27
|
+
nothing activates unconfirmed.
|
|
28
|
+
2. **The loop drives** — after every agent turn, glla nudges the work forward;
|
|
29
|
+
the widget (bottom-left) shows the goal, elapsed time, and last action.
|
|
30
|
+
3. **Verified completion** — when the agent calls `complete_goal`, glla
|
|
31
|
+
runs deterministic mechanical pre-audit checks first (~200ms fast-fail on
|
|
32
|
+
compiler/test failures) before queueing the **detached auditor worker process** (a
|
|
33
|
+
fresh pi RPC session with no extensions). It re-runs your checks and
|
|
34
|
+
demands raw output per contract item without holding the main pi turn.
|
|
35
|
+
Done sticks only when the auditor approves with evidence.
|
|
36
|
+
|
|
37
|
+
Then the other two modes:
|
|
38
|
+
|
|
39
|
+
```
|
|
40
|
+
/list add "refactor the cache layer" "add tests for the parser" # an audited queue
|
|
41
|
+
/loop audit # forever project-audit cadence
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
## What you should see
|
|
45
|
+
|
|
46
|
+
- **Commands**: `/goal`, `/list`, `/loop`, `/glla` (settings), `/review`.
|
|
47
|
+
- **Widget**: live goal/loop status — objective, elapsed time, last tool action.
|
|
48
|
+
- **State on disk**: `.pi-glla/` in your project — ledger `active.jsonl`,
|
|
49
|
+
goal markdown in `goals/`, finished goals in `archive/`.
|
|
50
|
+
- **Tools for the agent** (only while a goal is active): `complete_goal`,
|
|
51
|
+
`pause_goal`, `complete_task`, `update_task_status`.
|
|
52
|
+
|
|
53
|
+
## Install from source (developers)
|
|
54
|
+
|
|
55
|
+
Prerequisites: Node 22+ and bun (the test runner — `bun test`), pi-coding-agent, TypeScript 5.9+ (for `tsc --noEmit`).
|
|
56
|
+
|
|
57
|
+
```bash
|
|
58
|
+
git clone https://github.com/DraconDev/pi-goal-list-loop-audit.git # or use the local dir
|
|
59
|
+
cd pi-goal-list-loop-audit
|
|
60
|
+
pi install . # installs from local path
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
## Try it without installing
|
|
64
|
+
|
|
65
|
+
```bash
|
|
66
|
+
pi -e /home/dracon/Dev/pi-goal-list-loop-audit
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
---
|
|
70
|
+
|
|
71
|
+
## Operator notes (by release)
|
|
72
|
+
|
|
73
|
+
Everything below is reference material written as features landed — you don't
|
|
74
|
+
need it to get started.
|
|
75
|
+
|
|
76
|
+
## Auditor model: the built-in-provider rule
|
|
77
|
+
|
|
78
|
+
The auditor runs in a **detached fresh pi RPC process with no extensions**, so it can only use **built-in providers** (opencode, openrouter, minimax, google, anthropic, …).
|
|
79
|
+
You select the model in pi; the auditor uses it. The plugin never picks a
|
|
80
|
+
model itself. The resolution is just:
|
|
81
|
+
|
|
82
|
+
1. your explicit auditor pick — `/glla → Auditor model`, else
|
|
83
|
+
2. the pi session model — whatever you selected in pi.
|
|
84
|
+
|
|
85
|
+
If your session model's provider is extension-registered, the auditor's
|
|
86
|
+
extension-less session cannot auth it and the plugin says so at session start,
|
|
87
|
+
with the two fixes: switch pi's model to a built-in provider, or choose a
|
|
88
|
+
working auditor model in `/glla → Auditor model`.
|
|
89
|
+
|
|
90
|
+
Whatever you choose must work extension-less. Verify with:
|
|
91
|
+
|
|
92
|
+
```bash
|
|
93
|
+
PI_CODING_AGENT_DIR=/tmp/bare-agent pi -p "say ok" --model "provider/model-id"
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
The detached worker resolves `pi` from `PATH` and inherits the normal
|
|
97
|
+
`PI_CODING_AGENT_DIR`/provider environment. Set `GLLA_PI_BINARY=/absolute/path/to/pi`
|
|
98
|
+
when the CLI is not on the worker's PATH; credentials are never written into
|
|
99
|
+
`.pi-glla/audit-jobs/` or command arguments.
|
|
100
|
+
|
|
101
|
+
## Loop behavior: the multi-signal stuck gate (v0.25.1)
|
|
102
|
+
|
|
103
|
+
A `/loop` iteration is judged STUCK only when **every** progress signal is
|
|
104
|
+
zero — no file writes (`write`/`edit`/`multi_edit`/`write_file` tool
|
|
105
|
+
results), no git commits since the iteration began (HEAD advance), no
|
|
106
|
+
`spec_item_progress` ledger events, and no *paired* forward transition
|
|
107
|
+
("Next step (iter-N…)" text only counts when the same iteration also wrote
|
|
108
|
+
a file or committed — narration alone is the narrate-but-don't-ship loop)
|
|
109
|
+
— **and** the legacy same-tool-same-result check also fires.
|
|
110
|
+
|
|
111
|
+
Why it changed: the v0.24.0 single-signal detector (same tool + same
|
|
112
|
+
result hash 3×) killed two real user loops that were shipping work with
|
|
113
|
+
stable verification output — stable verification is the GOAL state of a
|
|
114
|
+
metricless loop, not the stuck state. Design doc:
|
|
115
|
+
`audit/STUCK-DETECTION-REWORK-2026-07-24.md`. `/loop start toolsamerepeat=0`
|
|
116
|
+
disables the legacy check entirely; `/loop finish [reason]` ends a loop
|
|
117
|
+
cleanly with stopReason `completed: <reason>` (distinct from
|
|
118
|
+
stuck/plateau/stopped-by-user).
|
|
119
|
+
|
|
120
|
+
## Provider recovery + aggressive mode
|
|
121
|
+
|
|
122
|
+
**Reason-agnostic retry.** Provider wording and upstream retry hints are not
|
|
123
|
+
used as availability or quota checks. Any retriable auditor failure is
|
|
124
|
+
infrastructure, not a verdict: the goal pauses with one eager 5-second retry,
|
|
125
|
+
then retries at the next `:00:30` slot after each hour starts. The durable
|
|
126
|
+
attempt and 24-hour bounds prevent an unbounded worker storm. `/goal resume`
|
|
127
|
+
retries immediately; a user pause is never stomped. Main-model recovery uses
|
|
128
|
+
the same generic policy and an ordered backup chain when configured.
|
|
129
|
+
|
|
130
|
+
**Aggressive mode** (Settings → Aggressive mode in `/glla`) flips the
|
|
131
|
+
continuation DEFAULTS toward keep-going:
|
|
132
|
+
|
|
133
|
+
| Key | default | aggressive |
|
|
134
|
+
|---|---|---|
|
|
135
|
+
| autoResume | default (hold on session load) | on (GLOBAL-only since v0.29.5 — project keys inert) |
|
|
136
|
+
| auditCap | 5 | 10 |
|
|
137
|
+
| stuckMaxInterventions | 5 | 10 |
|
|
138
|
+
| wedgeAlertMinutes | 30 | 0 (off) |
|
|
139
|
+
|
|
140
|
+
Explicit per-key settings always win — aggressiveMode flips defaults, never
|
|
141
|
+
your choices. Under aggressive mode an audit-cap disapproval streak does
|
|
142
|
+
NOT pause: the auditor's objections become a TODO list (`pendingTasks`)
|
|
143
|
+
rendered into every continuation, and the goal stays ACTIVE. Every
|
|
144
|
+
auto-event announces itself with a one-line notify.
|
|
145
|
+
|
|
146
|
+
## Subagent model inheritance (v0.24.6)
|
|
147
|
+
|
|
148
|
+
If you use `@tintinweb/pi-subagents`: its default `Explore` agent pins
|
|
149
|
+
`anthropic/claude-haiku-4-5`, so `Explore` subagents run on a **different
|
|
150
|
+
provider and quota pool than your session** — a quota-capped key (e.g.
|
|
151
|
+
OpenRouter) 403s after a few concurrent spawns even while the parent
|
|
152
|
+
session is fine.
|
|
153
|
+
|
|
154
|
+
glla fixes this by default: at session start it manages
|
|
155
|
+
`~/.pi/agent/agents/Explore.md` (pi-subagents' native override mechanism)
|
|
156
|
+
without the model pin, so subagents inherit your session model. Your own
|
|
157
|
+
same-named files are never touched (glla only edits files carrying its
|
|
158
|
+
`x-managed-by` marker).
|
|
159
|
+
|
|
160
|
+
Control it via `/glla` → Settings:
|
|
161
|
+
|
|
162
|
+
- **Subagent model strategy** — `inherit-parent` (default, subagents share
|
|
163
|
+
your session model + quota) or `agent-default` (upstream: Explore pins
|
|
164
|
+
haiku — cheap search, separate quota).
|
|
165
|
+
- **Subagent Explore model pin** — e.g. `minimax/MiniMax-M3`; always wins
|
|
166
|
+
over strategy.
|
|
167
|
+
|
|
168
|
+
Changes apply to NEW pi sessions (pi-subagents registers agents at its own
|
|
169
|
+
session start).
|
|
170
|
+
|
|
171
|
+
Release-workflow note: installing into the local extension tree
|
|
172
|
+
(`~/.pi/agent/npm`) requires `--legacy-peer-deps` — a pre-existing
|
|
173
|
+
`@pi-unipi/notify` peer pin on `@earendil-works/pi-coding-agent@^0.78.0`
|
|
174
|
+
conflicts with the current pi release.
|
|
175
|
+
|
|
176
|
+
## Run the tests
|
|
177
|
+
|
|
178
|
+
```bash
|
|
179
|
+
npm test
|
|
180
|
+
```
|
|
181
|
+
|
|
182
|
+
Expected output at v0.35.3: 1363 passing tests across 113 files (1 env-gated skip).
|
|
183
|
+
Counts change as bounded regressions are added; use the command output
|
|
184
|
+
as the source of truth ("N pass / 0 fail").
|
|
185
|
+
|
|
186
|
+
## Run the type-check
|
|
187
|
+
|
|
188
|
+
```bash
|
|
189
|
+
npm run check
|
|
190
|
+
```
|
|
191
|
+
|
|
192
|
+
Expected output: no TypeScript errors.
|
|
193
|
+
|
|
194
|
+
## End-to-end smoke test
|
|
195
|
+
|
|
196
|
+
After installing:
|
|
197
|
+
|
|
198
|
+
1. In a pi session, run:
|
|
199
|
+
```
|
|
200
|
+
/goal start "
|
|
201
|
+
Add a /healthz endpoint to src/server.ts that returns {status:'ok'} JSON.
|
|
202
|
+
|
|
203
|
+
Done when:
|
|
204
|
+
- curl -fsS localhost:3000/healthz returns 200 with body {\"status\":\"ok\"}
|
|
205
|
+
- The file is committed
|
|
206
|
+
"
|
|
207
|
+
```
|
|
208
|
+
2. The orchestrator creates `.pi-glla/goals/<id>.md`, schedules continuation, and the agent starts.
|
|
209
|
+
3. The agent reads the goal, makes the change, runs the verification, and calls `complete_goal`.
|
|
210
|
+
4. The orchestrator queues a detached auditor worker and returns control to the main turn.
|
|
211
|
+
5. The worker inspects files, runs `curl`, reads `git log`, and writes an identity-checked result.
|
|
212
|
+
6. Either `<approved/>` → goal archived; or `<disapproved/>` → loop continues.
|
|
213
|
+
|
|
214
|
+
## Reading the state
|
|
215
|
+
|
|
216
|
+
While the loop runs:
|
|
217
|
+
|
|
218
|
+
```bash
|
|
219
|
+
ls .pi-glla/ # see live state
|
|
220
|
+
cat .pi-glla/active.jsonl | tail -5
|
|
221
|
+
cat .pi-glla/goals/<id>.md # current goal markdown
|
|
222
|
+
ls .pi-glla/archive # past goals
|
|
223
|
+
```
|
|
224
|
+
|
|
225
|
+
## v0.1.0 verification status (2026-07-20, all live-verified)
|
|
226
|
+
|
|
227
|
+
- [x] Live `agent_end` loop fires after agent returns.
|
|
228
|
+
- [x] `complete_goal` triggers the isolated auditor session.
|
|
229
|
+
- [x] Auditor session correctly isolates (no extensions — discovered the built-in-provider rule).
|
|
230
|
+
- [x] `<approved/>` archives the goal with clean history.
|
|
231
|
+
- [x] `<disapproved/>` / auditor error continues or pauses with feedback.
|
|
232
|
+
- [x] 5-consecutive-error auto-pause fires (verified via live 403 storm).
|
|
233
|
+
- [x] Stale-ctx safety after session replacement (lastCtx pattern).
|
|
234
|
+
- [x] `npm test` 24/24. `npm run check` clean.
|
|
235
|
+
|
|
236
|
+
Current behavior: completion audits are detached from the main pi process.
|
|
237
|
+
Use `/goal status` to distinguish queued/running/recovery-pending work and
|
|
238
|
+
`/goal cancel` to discard a pending claim; a fresh `/goal resume` starts a new
|
|
239
|
+
attempt after lifecycle replacement.
|
|
240
|
+
|
|
241
|
+
## Reading your glla telemetry (v0.25.2)
|
|
242
|
+
|
|
243
|
+
`/glla stats` scans `.pi-glla/active.jsonl` across every project on the
|
|
244
|
+
rig and prints a per-project rollup: goals created, audit verdicts
|
|
245
|
+
(approved / disapproved / infra errors), average turns and file writes per
|
|
246
|
+
goal, premature-success count, total tokens, and last activity.
|
|
247
|
+
|
|
248
|
+
**Premature success** = an approved goal with < 50 turns AND < 5 file
|
|
249
|
+
writes AND < 8 bash calls — the "claimed done in 12 turns with 0 file
|
|
250
|
+
writes" pattern an auditor should have caught. `/glla stats premature`
|
|
251
|
+
lists only those projects, worst ratio first. Goals archived before
|
|
252
|
+
v0.25.2 have no telemetry and are never flagged retroactively.
|
|
253
|
+
|
|
254
|
+
`/glla stats json` emits the same rows as JSON (pipe to `jq`);
|
|
255
|
+
`/glla stats project=~/Dev/xyz` scopes to one project. `total_cost` is
|
|
256
|
+
measured in tokens (this rig has no per-provider price table).
|
|
257
|
+
|
|
258
|
+
## Modes (v0.25.3)
|
|
259
|
+
|
|
260
|
+
The three loops are not redundant — each long-runs differently:
|
|
261
|
+
|
|
262
|
+
| Mode | Item size | Long-running by |
|
|
263
|
+
|---|---|---|
|
|
264
|
+
| `/goal` | ONE big multi-hour task | Scope |
|
|
265
|
+
| `/list` | N items × short (minutes each) | Queue depth |
|
|
266
|
+
| `/loop` | 1 metric × infinite polish | Bounds |
|
|
267
|
+
|
|
268
|
+
`/list` items should fit in a single agent run; hundreds of them in the
|
|
269
|
+
queue is the right framing. `/list depth` shows queue depth, oldest item
|
|
270
|
+
age, and average item duration. Drafting cross-recommends: multi-hour
|
|
271
|
+
seeds in `/list` get pointed at `/goal`, aggregate "N items, one commit
|
|
272
|
+
each" seeds get shaped into N short items. See **LIST-PHILOSOPHY.md**
|
|
273
|
+
for the full hierarchy and the wrapper-goal anti-pattern it prevents.
|
|
274
|
+
|
|
275
|
+
## Auditing the auditor (v0.25.4)
|
|
276
|
+
|
|
277
|
+
Every audit verdict is appended to `.pi-glla/audits.jsonl` (goal id,
|
|
278
|
+
verdict, model, full report) — the durable trail for "where are we weak"
|
|
279
|
+
reviews. `/glla audits` lists the last 10 verdicts, `/glla audits 30`
|
|
280
|
+
shows more, `/glla audits full` prints the latest report. Reports are
|
|
281
|
+
think-block-stripped; disapprovals end with a `## Required fixes`
|
|
282
|
+
actionable tail, which is also what capped executor feedback keeps.
|
|
283
|
+
|
|
284
|
+
## Reviewer (postaudit since v0.27.5) — post-completion follow-up enqueuer
|
|
285
|
+
|
|
286
|
+
When a `/goal` completes or a `/list` queue empties, the reviewer fires:
|
|
287
|
+
it reads the archive + audit reports, extracts findings, classifies them
|
|
288
|
+
by **leverage**, writes a report to `.pi-glla/reviews/<goal-id>-<ts>.md`,
|
|
289
|
+
and cascades:
|
|
290
|
+
|
|
291
|
+
| Finding class | Action | Confirm? |
|
|
292
|
+
|---|---|---|
|
|
293
|
+
| Bug (`TODO`, `FIXME`, `bug`, `regression`, `broken`) | `/list` items | No — fix-without-confirm |
|
|
294
|
+
| Refactor (`duplicated`, `could be cleaner`, `left out`) | `/list` items | No |
|
|
295
|
+
| Architectural (`rewrite`, `new dependency`, `schema change`) | `/goal` proposal | Yes |
|
|
296
|
+
| Strategic (`should we…`, `deprecate`) | notify only | — |
|
|
297
|
+
| Clean completion (no findings) | audit `/goal` proposal | Yes |
|
|
298
|
+
|
|
299
|
+
The leverage principle: if you'd never say no to fixing a bug, the
|
|
300
|
+
reviewer doesn't ask. Decisions stay with you.
|
|
301
|
+
|
|
302
|
+
**Modes** (`/glla postaudit` → Mode — `/glla reviewer` is a kept alias —
|
|
303
|
+
or `/review <id> <mode>` for a one-shot override):
|
|
304
|
+
|
|
305
|
+
| Mode | Problems / improvements found | Architectural | Clean completion |
|
|
306
|
+
|---|---|---|---|
|
|
307
|
+
| `off` | reviewer never fires | — | — |
|
|
308
|
+
| `on` (default) | `/list` items, no Confirm | `/goal` proposal (Confirm) | audit `/goal` proposal (Confirm) |
|
|
309
|
+
| `auto` | `/list` items, no Confirm | `/list` items, no Confirm | audit enqueued as a `/list` item, no Confirm |
|
|
310
|
+
| `aggressive` | `/list` items, no Confirm | `/list` items + the first finding **relaunched as the next active `/goal`** | the regression-scan audit **relaunched as `/goal`** directly |
|
|
311
|
+
|
|
312
|
+
(v0.27.9 replaced the old `default`/`report` modes with this 4-mode set:
|
|
313
|
+
`default` → `on`; `report` was dropped — a silent report with no cascade
|
|
314
|
+
was the do-nothing mode.)
|
|
315
|
+
|
|
316
|
+
`auto` is the **auto-loop**: run it once and the cascade keeps rolling
|
|
317
|
+
through everything it finds — problems, improvements ("consider
|
|
318
|
+
adding…", "could be improved", "enhancement" are extracted too), then
|
|
319
|
+
the regression-scan audit — until the findings run dry. `aggressive`
|
|
320
|
+
goes one step further: the queue is skipped for the headline item — the
|
|
321
|
+
first architectural finding (or the clean-completion audit) relaunches
|
|
322
|
+
as the next ACTIVE goal with no Confirm at all, so the unattended rig
|
|
323
|
+
never stops. Strategic
|
|
324
|
+
findings (`should we…`) stay notify-only in every mode: decisions never
|
|
325
|
+
auto-fire. Extraction ignores code lines, markdown tables, code spans, and the
|
|
326
|
+
reviewer's own report vocabulary (v0.26.3), and findings are mined only
|
|
327
|
+
from the archive plus DISAPPROVED/error audit reports — an approved
|
|
328
|
+
report is the executor's self-claims, zero finding signal (v0.26.4,
|
|
329
|
+
after a second live self-match on the 0.26.3 completion). Stalls are
|
|
330
|
+
watched three ways: refire streaks and a pending-latch watchdog (a queued
|
|
331
|
+
continuation whose turn trigger was dropped — seen post-compaction) both
|
|
332
|
+
escalate to a loud pause/stop, and busy-session wedges alert at 30m
|
|
333
|
+
(v0.26.5). The heartbeat never suppresses itself on "recent ship" — that
|
|
334
|
+
heuristic self-sustained via state-file mtime (v0.26.6, after a 9.1h
|
|
335
|
+
darklord stall). In `auto` the 5-minute refire window is skipped for
|
|
336
|
+
list-complete events (the queue emptying is the cascade's natural
|
|
337
|
+
rhythm); the per-day cap (`maxReviewsPerDay`, default 20) still bounds
|
|
338
|
+
everything.
|
|
339
|
+
|
|
340
|
+
Safety: no firing on aborts/pauses, a 5-minute refire window blocks
|
|
341
|
+
runaway recursion, `maxReviewsPerDay: 20` caps the day, and `/loop`
|
|
342
|
+
never triggers it. Configure per-project via `/glla postaudit`
|
|
343
|
+
(mode, triggers, cascade steps, caps) — the block lives
|
|
344
|
+
in `.pi-glla/settings.json` under `postaudit` (the legacy `reviewer` key
|
|
345
|
+
is still read). Re-review any archived goal with
|
|
346
|
+
`/review <goal-id>` (bypasses the trigger gates).
|
|
347
|
+
|
|
348
|
+
## Stall handling (v0.26.1) — the zombie killer
|
|
349
|
+
|
|
350
|
+
Motivating incident (hegemon, 2026-07-25/26): a metricless spec loop
|
|
351
|
+
stopped producing turns; the heartbeat re-fired every 60s for **23.5
|
|
352
|
+
hours** (619 refires, zero turns, zero tokens) while the status line
|
|
353
|
+
still read "active". Three gaps made it invisible: the send path was
|
|
354
|
+
silent, the nudge counter counts *turns* (a zombie runs none), and no
|
|
355
|
+
compaction hook existed.
|
|
356
|
+
|
|
357
|
+
What ships:
|
|
358
|
+
|
|
359
|
+
- **Send-path ledger instrumentation** — `loop_turn_sent` /
|
|
360
|
+
`loop_turn_send_failed` (with the error text) and
|
|
361
|
+
`goal_continuation_sent` / `goal_continuation_send_failed` are now in
|
|
362
|
+
`.pi-glla/active.jsonl`. A stall is diagnosable from the ledger alone:
|
|
363
|
+
refires without matching `*_sent` = the send is throwing; `*_sent`
|
|
364
|
+
without a following turn = the turn trigger is dead.
|
|
365
|
+
- **Refire-streak escalation** — consecutive heartbeat refires that
|
|
366
|
+
produce no real agent turn are counted (reset only by `agent_end` /
|
|
367
|
+
`tool_call`, never by the refire itself). At the threshold (default 5;
|
|
368
|
+
edit Stall escalation refires in `/glla`, 0 = never) the supervisor stops spinning:
|
|
369
|
+
the loop stops / the goal pauses with `stalled: continuation not
|
|
370
|
+
landing`, a `stall_escalated` ledger event, a TUI warning, and an
|
|
371
|
+
external notify. The fix on the box: restart pi, resume.
|
|
372
|
+
- **Compaction hook** — `session_compact` now re-arms the continuation
|
|
373
|
+
chain ~2s after compaction when the session is idle with nothing
|
|
374
|
+
scheduled (`session_compact` + `compaction_refire` ledger events), so
|
|
375
|
+
post-compaction recovery no longer waits for the 60s heartbeat.
|
|
376
|
+
- **Stall surface** — the status line and widget show `stalls:N` while
|
|
377
|
+
the streak is nonzero, so a spinning supervisor is visible at a glance.
|
package/README.md
CHANGED
|
@@ -10,7 +10,7 @@ This is a detached process, not a nested session in the main pi process. `comple
|
|
|
10
10
|
|
|
11
11
|
On Windows, npm installs the `pi.cmd` shim rather than a directly executable `pi` binary. The auditor launches it through an explicitly quoted `cmd.exe` boundary; POSIX keeps direct shell-less execution. Protocol snapshots also tolerate transient Windows file-locks without deleting the last valid snapshot first.
|
|
12
12
|
|
|
13
|
-
**Current package version:** `v0.35.
|
|
13
|
+
**Current package version:** `v0.35.13` — use `/glla version` to see the installed version and the command for comparing it with the registry latest. This checkout may contain unreleased changes; the npm registry is authoritative for published versions.
|
|
14
14
|
|
|
15
15
|
## Why this exists
|
|
16
16
|
|
|
@@ -22,9 +22,12 @@ Most pi goal extensions — `pi-goal`, `pi-goal-x`, `pi-loop-mode`, `ralphi`, `t
|
|
|
22
22
|
|
|
23
23
|
| Stage | Protection |
|
|
24
24
|
|---|---|
|
|
25
|
-
| Goal intake |
|
|
26
|
-
| Implementation | `agent_end`-driven
|
|
27
|
-
|
|
|
25
|
+
| Goal intake | Deep upfront grilling + Confirm/Reject dialog; nothing activates unconfirmed |
|
|
26
|
+
| Implementation | Zero-pause autonomous execution: `agent_end`-driven loop with autonomous pivoting & sensible defaults |
|
|
27
|
+
| Milestone gates | Structured task milestones with mechanical test execution before completion |
|
|
28
|
+
| Pre-Audit | **Deterministic Fast-Fail**: Mechanical shell checks (`npm test`, `tsc`, `cargo test`) run in ~200ms before spawning the auditor worker |
|
|
29
|
+
| Final Verification | Detached extension-less auditor process + **regression_shield**: raw command output required per verification-contract item, enforced orchestrator-side |
|
|
30
|
+
| Queue Hygiene | Anti-queue-drift: every single `/list` item receives an independent detached audit pass before queue advancement |
|
|
28
31
|
|
|
29
32
|
## Quick start
|
|
30
33
|
|
|
@@ -314,9 +317,14 @@ leaving the list). Every recoverable provider failure uses the same ordered
|
|
|
314
317
|
chain: glla calls `setModel` for the first eligible fallback, the next
|
|
315
318
|
supervised turn tests it, and later failures advance left-to-right. Forbidden,
|
|
316
319
|
unavailable, and unauthenticated references are skipped; a successful
|
|
317
|
-
supervised turn
|
|
318
|
-
|
|
319
|
-
|
|
320
|
+
supervised turn on a fallback proves only that fallback is healthy. By default
|
|
321
|
+
(`mainModelFailback=auto`), the original primary remains durable and is probed
|
|
322
|
+
again every `mainModelPrimaryProbeMinutes` (15 minutes by default); a successful
|
|
323
|
+
supervised primary turn fails back and clears the episode. Set
|
|
324
|
+
`mainModelFailback=sticky` to preserve the legacy stay-on-fallback behavior.
|
|
325
|
+
The chain is global, durable, and its attempted cursor survives reload. After
|
|
326
|
+
the chain is exhausted, bounded retries continue on the active model rather
|
|
327
|
+
than silently abandoning work.
|
|
320
328
|
The Main agent tab shows the `N/10` count and numbered chain. The Drafter and
|
|
321
329
|
Auditor tabs likewise show each selected model together with its requested
|
|
322
330
|
thinking level; fallback rows show the effective/requested thinking level when
|
|
@@ -436,7 +444,7 @@ Open `/glla` to edit these settings in the table (the rows show effective values
|
|
|
436
444
|
- Auditor fallback agent
|
|
437
445
|
- Notify command, token limit, and wedge-alert minutes
|
|
438
446
|
- Auto-resume, auto-accept drafts, decision popup, and carryover policy
|
|
439
|
-
- Main-agent current model/thinking, fallback models,
|
|
447
|
+
- Main-agent current model/thinking, fallback models, recovery cadence, and preferred-primary failback policy in the Main agent tab
|
|
440
448
|
- Drafter agent/thinking/fallback agents in the Drafter tab
|
|
441
449
|
- Auditor agent/thinking/fallback agent in the Auditor tab
|
|
442
450
|
- Forbidden model patterns and switch policy
|
|
@@ -452,8 +460,8 @@ There is no top-level `/glla key=value` setting syntax.
|
|
|
452
460
|
|
|
453
461
|
Resolution per key: **project > global > defaults** — EXCEPT `autoResume` and
|
|
454
462
|
agent recovery settings (`mainModelFallbacks`, `mainModelRetryMinutes`,
|
|
455
|
-
`
|
|
456
|
-
`hourlyRetryProbe`),
|
|
463
|
+
`mainModelFailback`, `mainModelPrimaryProbeMinutes`, `drafterModel`,
|
|
464
|
+
`drafterThinkingLevel`, `drafterModelFallbacks`, `hourlyRetryProbe`),
|
|
457
465
|
which are **global-only**: per-project opt-ins from old versions
|
|
458
466
|
silently overrode the global hold at launch (the junk-runner incident), so
|
|
459
467
|
the launch-restore gate and the reviewer-enqueue gate read only the global
|
|
@@ -467,9 +475,14 @@ unavailable, and unauthenticated refs are skipped. When every candidate is
|
|
|
467
475
|
down, glla stops the current send attempt and uses the configured
|
|
468
476
|
`base → 2×base → 4×base → 8×base → 16×base → 5h` ladder (`base` defaults to
|
|
469
477
|
15m). `hourlyRetryProbe=on` adds a blind :00:30 retry after each hour starts.
|
|
470
|
-
|
|
471
|
-
|
|
472
|
-
|
|
478
|
+
With `mainModelFailback=auto` (the default), a successful fallback keeps the
|
|
479
|
+
original primary as the preferred model and schedules a durable health probe at
|
|
480
|
+
the `mainModelPrimaryProbeMinutes` cadence; the primary is selected only for a
|
|
481
|
+
supervised probe, and a failure returns to the serving fallback. Set
|
|
482
|
+
`mainModelFailback=sticky` to disable this reverse probe. No provider
|
|
483
|
+
availability or quota check is made before any retry; all recoverable failures
|
|
484
|
+
walk the ordered fallbacks and then continue on the active model through the
|
|
485
|
+
bounded retry policy. Automatic recovery stops at 24h,
|
|
473
486
|
preserves the saved work, and requires an explicit
|
|
474
487
|
`/goal resume`, `/list resume`, or `/loop resume` to start a fresh window. A
|
|
475
488
|
provider becoming available within that horizon therefore resumes saved work
|
package/docs/DESIGN.md
CHANGED
|
@@ -220,10 +220,13 @@ architectural decisions that changed the SHAPE of the system:
|
|
|
220
220
|
provider hints win when in budget, while `hourlyQuotaProbe` is a separate
|
|
221
221
|
optional :00:30 ticker. A paused goal/held loop remains resumable. A fresh
|
|
222
222
|
startup obeys the existing `autoResume` consent gate.
|
|
223
|
-
- **Successful-turn reset**: a real non-error agent end
|
|
224
|
-
|
|
225
|
-
|
|
226
|
-
|
|
223
|
+
- **Successful-turn reset**: a real non-error agent end on the preferred
|
|
224
|
+
primary clears the recovery cycle. In the current failback policy, a
|
|
225
|
+
successful fallback turn instead arms the durable preferred-primary probe;
|
|
226
|
+
`mainModelFailback=sticky` retains the historical immediate reset. Manual
|
|
227
|
+
model selection cancels it; host restore selections do not, and user aborts
|
|
228
|
+
do not masquerade as success. Goal/list/loop cancellation clears its timer
|
|
229
|
+
and durable state.
|
|
227
230
|
|
|
228
231
|
## Addendum v0.34.48–v0.34.56 (lifecycle/recovery hardening — the stale-handle era)
|
|
229
232
|
|
|
@@ -337,6 +340,15 @@ replacement without delivering a successor `session_start`:
|
|
|
337
340
|
after every hour starts. Main-model recovery and detached-auditor recovery
|
|
338
341
|
use this same reason-agnostic rule, with existing context/user-abort and
|
|
339
342
|
safety-horizon exceptions.
|
|
343
|
+
- **Preferred-primary failback is durable and supervised**:
|
|
344
|
+
`mainModelFailback=auto` (the default) does not treat a successful fallback
|
|
345
|
+
turn as proof that the original primary is healthy. The recovery record keeps
|
|
346
|
+
`primary`, records `primaryProbeAt`, and uses
|
|
347
|
+
`mainModelPrimaryProbeMinutes` (15 by default) to select the primary for one
|
|
348
|
+
real supervised probe. A primary success clears the episode; a provider
|
|
349
|
+
failure walks back to the serving fallback and schedules the next reverse
|
|
350
|
+
probe. `sticky` preserves the legacy permanent fallback choice. The
|
|
351
|
+
`primaryProbeInFlight` marker and pending switch survive a session boundary.
|
|
340
352
|
- **Legacy state is inert**: old `quota-waiting` phases, quota-named retry
|
|
341
353
|
counters, and provider-hint fields are accepted only long enough to load
|
|
342
354
|
and normalize old files. Canonical persisted state uses `retry-waiting`,
|
|
@@ -539,9 +551,24 @@ This protects against model-generated summaries losing fidelity.
|
|
|
539
551
|
| LOW | Telegram push | v0.3.0 |
|
|
540
552
|
| LOW | Sub-task auto-close | v0.3.0 |
|
|
541
553
|
|
|
554
|
+
## Addendum v0.35.6 (Unattended autonomy, audit cadence, and parallelization)
|
|
555
|
+
|
|
556
|
+
- **Two-phase decision architecture (upfront grilling → zero pauses during execution)**:
|
|
557
|
+
- Drafting upfront is the sole interview boundary: the agent asks sharp questions about architecture, scope, error conditions, and test commands.
|
|
558
|
+
- Active execution is 100% unattended: the agent picks sensible architectural defaults, records rationale, and continues without interrupting the user for obvious choices or secondary questions. Non-blocking notes are deferred to the completion summary.
|
|
559
|
+
- Premium engineering standards: mandatory root-cause fixes, full TypeScript type safety, and comprehensive test coverage. If an approach fails verification after 2 attempts, the agent autonomously steps back and pivots to an alternative architecture.
|
|
560
|
+
- **Audit cadence across modes**:
|
|
561
|
+
- `/goal`: Evaluated by the detached isolated auditor at goal completion (`complete_goal`).
|
|
562
|
+
- `/list`: Evaluated by the detached isolated auditor at the completion of **every individual list task** before unlocking and activating item $N+1$. This prevents list drift, ensuring that errors in early tasks do not cascade into downstream tasks.
|
|
563
|
+
- `/loop`: Shell metric command evaluated on every iteration; LLM auditor does not run on intermediate iterations.
|
|
564
|
+
- **Parallelization architecture & opportunities**:
|
|
565
|
+
- *Current*: Parallel subagent exploration fan-out (spawning multiple `Explore` agents in one turn) and parallel disjoint-worktree implementation (`general-purpose` workers with `isolation: "worktree"`). Detached auditor runs concurrently in background OS process.
|
|
566
|
+
- *Opportunities*: Concurrent `/list` dispatch across non-conflicting tasks using isolated worktrees to eliminate head-of-line blocking for large queues, plus background contract rehearsals.
|
|
567
|
+
|
|
542
568
|
## Files
|
|
543
569
|
|
|
544
570
|
- `docs/DESIGN.md` — **this file**
|
|
545
571
|
- `README.md` — quickstart
|
|
546
572
|
- `audit/pi-name-v3-registry-based.md` — naming rationale
|
|
547
573
|
- `audit/pi-goal-loop-design.md` — earlier design (now superseded)
|
|
574
|
+
|
package/docs/INDEX.md
CHANGED
|
@@ -2,11 +2,34 @@
|
|
|
2
2
|
|
|
3
3
|
Ordered by reading path, not alphabetically.
|
|
4
4
|
|
|
5
|
+
## Active focus (recent work, durable artifacts)
|
|
6
|
+
|
|
7
|
+
This package's policy contracts and recent changes are recorded in
|
|
8
|
+
the audit/ directory of the **repository checkout** — it is not
|
|
9
|
+
shipped in the npm tarball (see "Repository-only material" below).
|
|
10
|
+
|
|
11
|
+
For shipped docs, the relevant entry points are:
|
|
12
|
+
|
|
13
|
+
- `../CHANGELOG.md` — user-facing changelog; the top of the file is the
|
|
14
|
+
current package version. v0.35.5 adopted the six-label completion
|
|
15
|
+
recap; v0.35.6 added typed-boundary regression pins; v0.35.7 added
|
|
16
|
+
deterministic fast-fail pre-audits, zero-pause autonomous execution, and
|
|
17
|
+
task milestone gating; v0.35.8 added main-model preferred-primary
|
|
18
|
+
failback; v0.35.9 hardened cross-version npm tarball checks; v0.35.10
|
|
19
|
+
handles multi-entry npm dry-run reports; v0.35.11 accepts both npm report
|
|
20
|
+
shapes; v0.35.12 supports npm 12's keyed pack reports; v0.35.13 fixes stale-API recovery loops.
|
|
21
|
+
- `../README.md` — what the plugin is, install, quickstart, and the
|
|
22
|
+
architectural guarantee (drafting + confirm + detached auditor).
|
|
23
|
+
- `../INSTALL.md` — manual install / symlink setup; the recommended
|
|
24
|
+
companion plugins and the `auditor reads / writes are path-checked`
|
|
25
|
+
note.
|
|
26
|
+
|
|
5
27
|
## Entry points
|
|
6
28
|
- `../README.md` — what the plugin is, install, quickstart
|
|
7
29
|
- `../INSTALL.md` — manual install / symlink setup
|
|
8
|
-
- `../
|
|
9
|
-
|
|
30
|
+
- `../CHANGELOG.md` — user-facing changelog; current package version is
|
|
31
|
+
at the top of the file (use `/glla version` to compare with the
|
|
32
|
+
registry).
|
|
10
33
|
|
|
11
34
|
## Architecture
|
|
12
35
|
- `DESIGN.md` — plugin design (types, state, extension lifecycle)
|
|
@@ -21,11 +44,13 @@ Ordered by reading path, not alphabetically.
|
|
|
21
44
|
- `../prompts/` — goal/loop drafting prompt templates
|
|
22
45
|
- `../schemas/` — goal state JSON schema
|
|
23
46
|
- `../examples/` — example objective files
|
|
24
|
-
- `../audit/INDEX.md` — the audit trail (every shipped change, newest first)
|
|
25
47
|
- `../CHANGELOG.md` — user-facing changelog (unreleased at top)
|
|
26
48
|
|
|
49
|
+
## Repository-only material
|
|
50
|
+
The audit history and competitor research live in `audit/` and `.research/`
|
|
51
|
+
for contributors, but are intentionally not included in the npm tarball.
|
|
52
|
+
|
|
27
53
|
## Research material
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
positioning doc's appendix for the package list.
|
|
54
|
+
`.research/` — competitor plugin sources pulled from npm tarballs for study
|
|
55
|
+
(gitignored, local only). Re-pull with `cd .research && npm pack <pkg> &&
|
|
56
|
+
tar xzf <tgz>`; see the positioning doc's appendix for the package list.
|