pi-goal-list-loop-audit 0.35.4 → 0.35.13

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/INSTALL.md ADDED
@@ -0,0 +1,377 @@
1
+ # Install & try
2
+
3
+ ## Install (the usual way)
4
+
5
+ ```bash
6
+ pi install npm:pi-goal-list-loop-audit
7
+ ```
8
+
9
+ That's it — pi loads the extension into every session (run `/reload` in any
10
+ session that was already open).
11
+
12
+ > **Persistence note**: `pi update` can overwrite `~/.pi/agent/npm/node_modules/`.
13
+ > If the plugin disappears after an update, re-run `pi install`. For a permanent
14
+ > install, copy the package into your project's `.pi/extensions/` directory instead.
15
+
16
+ ## Your first goal in 60 seconds
17
+
18
+ In any pi session, in the project you want worked on:
19
+
20
+ ```
21
+ /goal "Fix the login timeout bug. Done when: `npm test` passes"
22
+ ```
23
+
24
+ What happens:
25
+
26
+ 1. **Draft + Confirm** — glla shows the goal contract (objective + Done-when);
27
+ nothing activates unconfirmed.
28
+ 2. **The loop drives** — after every agent turn, glla nudges the work forward;
29
+ the widget (bottom-left) shows the goal, elapsed time, and last action.
30
+ 3. **Verified completion** — when the agent calls `complete_goal`, glla
31
+ runs deterministic mechanical pre-audit checks first (~200ms fast-fail on
32
+ compiler/test failures) before queueing the **detached auditor worker process** (a
33
+ fresh pi RPC session with no extensions). It re-runs your checks and
34
+ demands raw output per contract item without holding the main pi turn.
35
+ Done sticks only when the auditor approves with evidence.
36
+
37
+ Then the other two modes:
38
+
39
+ ```
40
+ /list add "refactor the cache layer" "add tests for the parser" # an audited queue
41
+ /loop audit # forever project-audit cadence
42
+ ```
43
+
44
+ ## What you should see
45
+
46
+ - **Commands**: `/goal`, `/list`, `/loop`, `/glla` (settings), `/review`.
47
+ - **Widget**: live goal/loop status — objective, elapsed time, last tool action.
48
+ - **State on disk**: `.pi-glla/` in your project — ledger `active.jsonl`,
49
+ goal markdown in `goals/`, finished goals in `archive/`.
50
+ - **Tools for the agent** (only while a goal is active): `complete_goal`,
51
+ `pause_goal`, `complete_task`, `update_task_status`.
52
+
53
+ ## Install from source (developers)
54
+
55
+ Prerequisites: Node 22+ and bun (the test runner — `bun test`), pi-coding-agent, TypeScript 5.9+ (for `tsc --noEmit`).
56
+
57
+ ```bash
58
+ git clone https://github.com/DraconDev/pi-goal-list-loop-audit.git # or use the local dir
59
+ cd pi-goal-list-loop-audit
60
+ pi install . # installs from local path
61
+ ```
62
+
63
+ ## Try it without installing
64
+
65
+ ```bash
66
+ pi -e /home/dracon/Dev/pi-goal-list-loop-audit
67
+ ```
68
+
69
+ ---
70
+
71
+ ## Operator notes (by release)
72
+
73
+ Everything below is reference material written as features landed — you don't
74
+ need it to get started.
75
+
76
+ ## Auditor model: the built-in-provider rule
77
+
78
+ The auditor runs in a **detached fresh pi RPC process with no extensions**, so it can only use **built-in providers** (opencode, openrouter, minimax, google, anthropic, …).
79
+ You select the model in pi; the auditor uses it. The plugin never picks a
80
+ model itself. The resolution is just:
81
+
82
+ 1. your explicit auditor pick — `/glla → Auditor model`, else
83
+ 2. the pi session model — whatever you selected in pi.
84
+
85
+ If your session model's provider is extension-registered, the auditor's
86
+ extension-less session cannot auth it and the plugin says so at session start,
87
+ with the two fixes: switch pi's model to a built-in provider, or choose a
88
+ working auditor model in `/glla → Auditor model`.
89
+
90
+ Whatever you choose must work extension-less. Verify with:
91
+
92
+ ```bash
93
+ PI_CODING_AGENT_DIR=/tmp/bare-agent pi -p "say ok" --model "provider/model-id"
94
+ ```
95
+
96
+ The detached worker resolves `pi` from `PATH` and inherits the normal
97
+ `PI_CODING_AGENT_DIR`/provider environment. Set `GLLA_PI_BINARY=/absolute/path/to/pi`
98
+ when the CLI is not on the worker's PATH; credentials are never written into
99
+ `.pi-glla/audit-jobs/` or command arguments.
100
+
101
+ ## Loop behavior: the multi-signal stuck gate (v0.25.1)
102
+
103
+ A `/loop` iteration is judged STUCK only when **every** progress signal is
104
+ zero — no file writes (`write`/`edit`/`multi_edit`/`write_file` tool
105
+ results), no git commits since the iteration began (HEAD advance), no
106
+ `spec_item_progress` ledger events, and no *paired* forward transition
107
+ ("Next step (iter-N…)" text only counts when the same iteration also wrote
108
+ a file or committed — narration alone is the narrate-but-don't-ship loop)
109
+ — **and** the legacy same-tool-same-result check also fires.
110
+
111
+ Why it changed: the v0.24.0 single-signal detector (same tool + same
112
+ result hash 3×) killed two real user loops that were shipping work with
113
+ stable verification output — stable verification is the GOAL state of a
114
+ metricless loop, not the stuck state. Design doc:
115
+ `audit/STUCK-DETECTION-REWORK-2026-07-24.md`. `/loop start toolsamerepeat=0`
116
+ disables the legacy check entirely; `/loop finish [reason]` ends a loop
117
+ cleanly with stopReason `completed: <reason>` (distinct from
118
+ stuck/plateau/stopped-by-user).
119
+
120
+ ## Provider recovery + aggressive mode
121
+
122
+ **Reason-agnostic retry.** Provider wording and upstream retry hints are not
123
+ used as availability or quota checks. Any retriable auditor failure is
124
+ infrastructure, not a verdict: the goal pauses with one eager 5-second retry,
125
+ then retries at the next `:00:30` slot after each hour starts. The durable
126
+ attempt and 24-hour bounds prevent an unbounded worker storm. `/goal resume`
127
+ retries immediately; a user pause is never stomped. Main-model recovery uses
128
+ the same generic policy and an ordered backup chain when configured.
129
+
130
+ **Aggressive mode** (Settings → Aggressive mode in `/glla`) flips the
131
+ continuation DEFAULTS toward keep-going:
132
+
133
+ | Key | default | aggressive |
134
+ |---|---|---|
135
+ | autoResume | default (hold on session load) | on (GLOBAL-only since v0.29.5 — project keys inert) |
136
+ | auditCap | 5 | 10 |
137
+ | stuckMaxInterventions | 5 | 10 |
138
+ | wedgeAlertMinutes | 30 | 0 (off) |
139
+
140
+ Explicit per-key settings always win — aggressiveMode flips defaults, never
141
+ your choices. Under aggressive mode an audit-cap disapproval streak does
142
+ NOT pause: the auditor's objections become a TODO list (`pendingTasks`)
143
+ rendered into every continuation, and the goal stays ACTIVE. Every
144
+ auto-event announces itself with a one-line notify.
145
+
146
+ ## Subagent model inheritance (v0.24.6)
147
+
148
+ If you use `@tintinweb/pi-subagents`: its default `Explore` agent pins
149
+ `anthropic/claude-haiku-4-5`, so `Explore` subagents run on a **different
150
+ provider and quota pool than your session** — a quota-capped key (e.g.
151
+ OpenRouter) 403s after a few concurrent spawns even while the parent
152
+ session is fine.
153
+
154
+ glla fixes this by default: at session start it manages
155
+ `~/.pi/agent/agents/Explore.md` (pi-subagents' native override mechanism)
156
+ without the model pin, so subagents inherit your session model. Your own
157
+ same-named files are never touched (glla only edits files carrying its
158
+ `x-managed-by` marker).
159
+
160
+ Control it via `/glla` → Settings:
161
+
162
+ - **Subagent model strategy** — `inherit-parent` (default, subagents share
163
+ your session model + quota) or `agent-default` (upstream: Explore pins
164
+ haiku — cheap search, separate quota).
165
+ - **Subagent Explore model pin** — e.g. `minimax/MiniMax-M3`; always wins
166
+ over strategy.
167
+
168
+ Changes apply to NEW pi sessions (pi-subagents registers agents at its own
169
+ session start).
170
+
171
+ Release-workflow note: installing into the local extension tree
172
+ (`~/.pi/agent/npm`) requires `--legacy-peer-deps` — a pre-existing
173
+ `@pi-unipi/notify` peer pin on `@earendil-works/pi-coding-agent@^0.78.0`
174
+ conflicts with the current pi release.
175
+
176
+ ## Run the tests
177
+
178
+ ```bash
179
+ npm test
180
+ ```
181
+
182
+ Expected output at v0.35.3: 1363 passing tests across 113 files (1 env-gated skip).
183
+ Counts change as bounded regressions are added; use the command output
184
+ as the source of truth ("N pass / 0 fail").
185
+
186
+ ## Run the type-check
187
+
188
+ ```bash
189
+ npm run check
190
+ ```
191
+
192
+ Expected output: no TypeScript errors.
193
+
194
+ ## End-to-end smoke test
195
+
196
+ After installing:
197
+
198
+ 1. In a pi session, run:
199
+ ```
200
+ /goal start "
201
+ Add a /healthz endpoint to src/server.ts that returns {status:'ok'} JSON.
202
+
203
+ Done when:
204
+ - curl -fsS localhost:3000/healthz returns 200 with body {\"status\":\"ok\"}
205
+ - The file is committed
206
+ "
207
+ ```
208
+ 2. The orchestrator creates `.pi-glla/goals/<id>.md`, schedules continuation, and the agent starts.
209
+ 3. The agent reads the goal, makes the change, runs the verification, and calls `complete_goal`.
210
+ 4. The orchestrator queues a detached auditor worker and returns control to the main turn.
211
+ 5. The worker inspects files, runs `curl`, reads `git log`, and writes an identity-checked result.
212
+ 6. Either `<approved/>` → goal archived; or `<disapproved/>` → loop continues.
213
+
214
+ ## Reading the state
215
+
216
+ While the loop runs:
217
+
218
+ ```bash
219
+ ls .pi-glla/ # see live state
220
+ cat .pi-glla/active.jsonl | tail -5
221
+ cat .pi-glla/goals/<id>.md # current goal markdown
222
+ ls .pi-glla/archive # past goals
223
+ ```
224
+
225
+ ## v0.1.0 verification status (2026-07-20, all live-verified)
226
+
227
+ - [x] Live `agent_end` loop fires after agent returns.
228
+ - [x] `complete_goal` triggers the isolated auditor session.
229
+ - [x] Auditor session correctly isolates (no extensions — discovered the built-in-provider rule).
230
+ - [x] `<approved/>` archives the goal with clean history.
231
+ - [x] `<disapproved/>` / auditor error continues or pauses with feedback.
232
+ - [x] 5-consecutive-error auto-pause fires (verified via live 403 storm).
233
+ - [x] Stale-ctx safety after session replacement (lastCtx pattern).
234
+ - [x] `npm test` 24/24. `npm run check` clean.
235
+
236
+ Current behavior: completion audits are detached from the main pi process.
237
+ Use `/goal status` to distinguish queued/running/recovery-pending work and
238
+ `/goal cancel` to discard a pending claim; a fresh `/goal resume` starts a new
239
+ attempt after lifecycle replacement.
240
+
241
+ ## Reading your glla telemetry (v0.25.2)
242
+
243
+ `/glla stats` scans `.pi-glla/active.jsonl` across every project on the
244
+ rig and prints a per-project rollup: goals created, audit verdicts
245
+ (approved / disapproved / infra errors), average turns and file writes per
246
+ goal, premature-success count, total tokens, and last activity.
247
+
248
+ **Premature success** = an approved goal with < 50 turns AND < 5 file
249
+ writes AND < 8 bash calls — the "claimed done in 12 turns with 0 file
250
+ writes" pattern an auditor should have caught. `/glla stats premature`
251
+ lists only those projects, worst ratio first. Goals archived before
252
+ v0.25.2 have no telemetry and are never flagged retroactively.
253
+
254
+ `/glla stats json` emits the same rows as JSON (pipe to `jq`);
255
+ `/glla stats project=~/Dev/xyz` scopes to one project. `total_cost` is
256
+ measured in tokens (this rig has no per-provider price table).
257
+
258
+ ## Modes (v0.25.3)
259
+
260
+ The three loops are not redundant — each long-runs differently:
261
+
262
+ | Mode | Item size | Long-running by |
263
+ |---|---|---|
264
+ | `/goal` | ONE big multi-hour task | Scope |
265
+ | `/list` | N items × short (minutes each) | Queue depth |
266
+ | `/loop` | 1 metric × infinite polish | Bounds |
267
+
268
+ `/list` items should fit in a single agent run; hundreds of them in the
269
+ queue is the right framing. `/list depth` shows queue depth, oldest item
270
+ age, and average item duration. Drafting cross-recommends: multi-hour
271
+ seeds in `/list` get pointed at `/goal`, aggregate "N items, one commit
272
+ each" seeds get shaped into N short items. See **LIST-PHILOSOPHY.md**
273
+ for the full hierarchy and the wrapper-goal anti-pattern it prevents.
274
+
275
+ ## Auditing the auditor (v0.25.4)
276
+
277
+ Every audit verdict is appended to `.pi-glla/audits.jsonl` (goal id,
278
+ verdict, model, full report) — the durable trail for "where are we weak"
279
+ reviews. `/glla audits` lists the last 10 verdicts, `/glla audits 30`
280
+ shows more, `/glla audits full` prints the latest report. Reports are
281
+ think-block-stripped; disapprovals end with a `## Required fixes`
282
+ actionable tail, which is also what capped executor feedback keeps.
283
+
284
+ ## Reviewer (postaudit since v0.27.5) — post-completion follow-up enqueuer
285
+
286
+ When a `/goal` completes or a `/list` queue empties, the reviewer fires:
287
+ it reads the archive + audit reports, extracts findings, classifies them
288
+ by **leverage**, writes a report to `.pi-glla/reviews/<goal-id>-<ts>.md`,
289
+ and cascades:
290
+
291
+ | Finding class | Action | Confirm? |
292
+ |---|---|---|
293
+ | Bug (`TODO`, `FIXME`, `bug`, `regression`, `broken`) | `/list` items | No — fix-without-confirm |
294
+ | Refactor (`duplicated`, `could be cleaner`, `left out`) | `/list` items | No |
295
+ | Architectural (`rewrite`, `new dependency`, `schema change`) | `/goal` proposal | Yes |
296
+ | Strategic (`should we…`, `deprecate`) | notify only | — |
297
+ | Clean completion (no findings) | audit `/goal` proposal | Yes |
298
+
299
+ The leverage principle: if you'd never say no to fixing a bug, the
300
+ reviewer doesn't ask. Decisions stay with you.
301
+
302
+ **Modes** (`/glla postaudit` → Mode — `/glla reviewer` is a kept alias —
303
+ or `/review <id> <mode>` for a one-shot override):
304
+
305
+ | Mode | Problems / improvements found | Architectural | Clean completion |
306
+ |---|---|---|---|
307
+ | `off` | reviewer never fires | — | — |
308
+ | `on` (default) | `/list` items, no Confirm | `/goal` proposal (Confirm) | audit `/goal` proposal (Confirm) |
309
+ | `auto` | `/list` items, no Confirm | `/list` items, no Confirm | audit enqueued as a `/list` item, no Confirm |
310
+ | `aggressive` | `/list` items, no Confirm | `/list` items + the first finding **relaunched as the next active `/goal`** | the regression-scan audit **relaunched as `/goal`** directly |
311
+
312
+ (v0.27.9 replaced the old `default`/`report` modes with this 4-mode set:
313
+ `default` → `on`; `report` was dropped — a silent report with no cascade
314
+ was the do-nothing mode.)
315
+
316
+ `auto` is the **auto-loop**: run it once and the cascade keeps rolling
317
+ through everything it finds — problems, improvements ("consider
318
+ adding…", "could be improved", "enhancement" are extracted too), then
319
+ the regression-scan audit — until the findings run dry. `aggressive`
320
+ goes one step further: the queue is skipped for the headline item — the
321
+ first architectural finding (or the clean-completion audit) relaunches
322
+ as the next ACTIVE goal with no Confirm at all, so the unattended rig
323
+ never stops. Strategic
324
+ findings (`should we…`) stay notify-only in every mode: decisions never
325
+ auto-fire. Extraction ignores code lines, markdown tables, code spans, and the
326
+ reviewer's own report vocabulary (v0.26.3), and findings are mined only
327
+ from the archive plus DISAPPROVED/error audit reports — an approved
328
+ report is the executor's self-claims, zero finding signal (v0.26.4,
329
+ after a second live self-match on the 0.26.3 completion). Stalls are
330
+ watched three ways: refire streaks and a pending-latch watchdog (a queued
331
+ continuation whose turn trigger was dropped — seen post-compaction) both
332
+ escalate to a loud pause/stop, and busy-session wedges alert at 30m
333
+ (v0.26.5). The heartbeat never suppresses itself on "recent ship" — that
334
+ heuristic self-sustained via state-file mtime (v0.26.6, after a 9.1h
335
+ darklord stall). In `auto` the 5-minute refire window is skipped for
336
+ list-complete events (the queue emptying is the cascade's natural
337
+ rhythm); the per-day cap (`maxReviewsPerDay`, default 20) still bounds
338
+ everything.
339
+
340
+ Safety: no firing on aborts/pauses, a 5-minute refire window blocks
341
+ runaway recursion, `maxReviewsPerDay: 20` caps the day, and `/loop`
342
+ never triggers it. Configure per-project via `/glla postaudit`
343
+ (mode, triggers, cascade steps, caps) — the block lives
344
+ in `.pi-glla/settings.json` under `postaudit` (the legacy `reviewer` key
345
+ is still read). Re-review any archived goal with
346
+ `/review <goal-id>` (bypasses the trigger gates).
347
+
348
+ ## Stall handling (v0.26.1) — the zombie killer
349
+
350
+ Motivating incident (hegemon, 2026-07-25/26): a metricless spec loop
351
+ stopped producing turns; the heartbeat re-fired every 60s for **23.5
352
+ hours** (619 refires, zero turns, zero tokens) while the status line
353
+ still read "active". Three gaps made it invisible: the send path was
354
+ silent, the nudge counter counts *turns* (a zombie runs none), and no
355
+ compaction hook existed.
356
+
357
+ What ships:
358
+
359
+ - **Send-path ledger instrumentation** — `loop_turn_sent` /
360
+ `loop_turn_send_failed` (with the error text) and
361
+ `goal_continuation_sent` / `goal_continuation_send_failed` are now in
362
+ `.pi-glla/active.jsonl`. A stall is diagnosable from the ledger alone:
363
+ refires without matching `*_sent` = the send is throwing; `*_sent`
364
+ without a following turn = the turn trigger is dead.
365
+ - **Refire-streak escalation** — consecutive heartbeat refires that
366
+ produce no real agent turn are counted (reset only by `agent_end` /
367
+ `tool_call`, never by the refire itself). At the threshold (default 5;
368
+ edit Stall escalation refires in `/glla`, 0 = never) the supervisor stops spinning:
369
+ the loop stops / the goal pauses with `stalled: continuation not
370
+ landing`, a `stall_escalated` ledger event, a TUI warning, and an
371
+ external notify. The fix on the box: restart pi, resume.
372
+ - **Compaction hook** — `session_compact` now re-arms the continuation
373
+ chain ~2s after compaction when the session is idle with nothing
374
+ scheduled (`session_compact` + `compaction_refire` ledger events), so
375
+ post-compaction recovery no longer waits for the 60s heartbeat.
376
+ - **Stall surface** — the status line and widget show `stalls:N` while
377
+ the streak is nonzero, so a spinning supervisor is visible at a glance.
package/README.md CHANGED
@@ -10,7 +10,7 @@ This is a detached process, not a nested session in the main pi process. `comple
10
10
 
11
11
  On Windows, npm installs the `pi.cmd` shim rather than a directly executable `pi` binary. The auditor launches it through an explicitly quoted `cmd.exe` boundary; POSIX keeps direct shell-less execution. Protocol snapshots also tolerate transient Windows file-locks without deleting the last valid snapshot first.
12
12
 
13
- **Current package version:** `v0.35.3` — use `/glla version` to see the installed version and the command for comparing it with the registry latest. This checkout may contain unreleased changes; the npm registry is authoritative for published versions.
13
+ **Current package version:** `v0.35.13` — use `/glla version` to see the installed version and the command for comparing it with the registry latest. This checkout may contain unreleased changes; the npm registry is authoritative for published versions.
14
14
 
15
15
  ## Why this exists
16
16
 
@@ -22,9 +22,12 @@ Most pi goal extensions — `pi-goal`, `pi-goal-x`, `pi-loop-mode`, `ralphi`, `t
22
22
 
23
23
  | Stage | Protection |
24
24
  |---|---|
25
- | Goal intake | Drafting + Confirm/Reject dialog; nothing activates unconfirmed |
26
- | Implementation | `agent_end`-driven continuation loop with 5-minute hard backoff cap |
27
- | Completion | Detached extension-less auditor process + **regression_shield**: raw command output required per verification-contract item, enforced orchestrator-side |
25
+ | Goal intake | Deep upfront grilling + Confirm/Reject dialog; nothing activates unconfirmed |
26
+ | Implementation | Zero-pause autonomous execution: `agent_end`-driven loop with autonomous pivoting & sensible defaults |
27
+ | Milestone gates | Structured task milestones with mechanical test execution before completion |
28
+ | Pre-Audit | **Deterministic Fast-Fail**: Mechanical shell checks (`npm test`, `tsc`, `cargo test`) run in ~200ms before spawning the auditor worker |
29
+ | Final Verification | Detached extension-less auditor process + **regression_shield**: raw command output required per verification-contract item, enforced orchestrator-side |
30
+ | Queue Hygiene | Anti-queue-drift: every single `/list` item receives an independent detached audit pass before queue advancement |
28
31
 
29
32
  ## Quick start
30
33
 
@@ -314,9 +317,14 @@ leaving the list). Every recoverable provider failure uses the same ordered
314
317
  chain: glla calls `setModel` for the first eligible fallback, the next
315
318
  supervised turn tests it, and later failures advance left-to-right. Forbidden,
316
319
  unavailable, and unauthenticated references are skipped; a successful
317
- supervised turn clears the episode. The chain is global, durable, and its
318
- attempted cursor survives reload. After the chain is exhausted, bounded
319
- retries continue on the active model rather than silently abandoning work.
320
+ supervised turn on a fallback proves only that fallback is healthy. By default
321
+ (`mainModelFailback=auto`), the original primary remains durable and is probed
322
+ again every `mainModelPrimaryProbeMinutes` (15 minutes by default); a successful
323
+ supervised primary turn fails back and clears the episode. Set
324
+ `mainModelFailback=sticky` to preserve the legacy stay-on-fallback behavior.
325
+ The chain is global, durable, and its attempted cursor survives reload. After
326
+ the chain is exhausted, bounded retries continue on the active model rather
327
+ than silently abandoning work.
320
328
  The Main agent tab shows the `N/10` count and numbered chain. The Drafter and
321
329
  Auditor tabs likewise show each selected model together with its requested
322
330
  thinking level; fallback rows show the effective/requested thinking level when
@@ -436,7 +444,7 @@ Open `/glla` to edit these settings in the table (the rows show effective values
436
444
  - Auditor fallback agent
437
445
  - Notify command, token limit, and wedge-alert minutes
438
446
  - Auto-resume, auto-accept drafts, decision popup, and carryover policy
439
- - Main-agent current model/thinking, fallback models, and recovery cadence in the Main agent tab
447
+ - Main-agent current model/thinking, fallback models, recovery cadence, and preferred-primary failback policy in the Main agent tab
440
448
  - Drafter agent/thinking/fallback agents in the Drafter tab
441
449
  - Auditor agent/thinking/fallback agent in the Auditor tab
442
450
  - Forbidden model patterns and switch policy
@@ -452,8 +460,8 @@ There is no top-level `/glla key=value` setting syntax.
452
460
 
453
461
  Resolution per key: **project > global > defaults** — EXCEPT `autoResume` and
454
462
  agent recovery settings (`mainModelFallbacks`, `mainModelRetryMinutes`,
455
- `drafterModel`, `drafterThinkingLevel`, `drafterModelFallbacks`,
456
- `hourlyRetryProbe`),
463
+ `mainModelFailback`, `mainModelPrimaryProbeMinutes`, `drafterModel`,
464
+ `drafterThinkingLevel`, `drafterModelFallbacks`, `hourlyRetryProbe`),
457
465
  which are **global-only**: per-project opt-ins from old versions
458
466
  silently overrode the global hold at launch (the junk-runner incident), so
459
467
  the launch-restore gate and the reviewer-enqueue gate read only the global
@@ -467,9 +475,14 @@ unavailable, and unauthenticated refs are skipped. When every candidate is
467
475
  down, glla stops the current send attempt and uses the configured
468
476
  `base → 2×base → 4×base → 8×base → 16×base → 5h` ladder (`base` defaults to
469
477
  15m). `hourlyRetryProbe=on` adds a blind :00:30 retry after each hour starts.
470
- No provider availability or quota check is made before any retry; all
471
- recoverable failures walk the ordered fallbacks and then continue on the active
472
- model through the bounded retry policy. Automatic recovery stops at 24h,
478
+ With `mainModelFailback=auto` (the default), a successful fallback keeps the
479
+ original primary as the preferred model and schedules a durable health probe at
480
+ the `mainModelPrimaryProbeMinutes` cadence; the primary is selected only for a
481
+ supervised probe, and a failure returns to the serving fallback. Set
482
+ `mainModelFailback=sticky` to disable this reverse probe. No provider
483
+ availability or quota check is made before any retry; all recoverable failures
484
+ walk the ordered fallbacks and then continue on the active model through the
485
+ bounded retry policy. Automatic recovery stops at 24h,
473
486
  preserves the saved work, and requires an explicit
474
487
  `/goal resume`, `/list resume`, or `/loop resume` to start a fresh window. A
475
488
  provider becoming available within that horizon therefore resumes saved work
package/docs/DESIGN.md CHANGED
@@ -220,10 +220,13 @@ architectural decisions that changed the SHAPE of the system:
220
220
  provider hints win when in budget, while `hourlyQuotaProbe` is a separate
221
221
  optional :00:30 ticker. A paused goal/held loop remains resumable. A fresh
222
222
  startup obeys the existing `autoResume` consent gate.
223
- - **Successful-turn reset**: a real non-error agent end clears the recovery
224
- cycle. Manual model selection cancels it; host restore selections do not, and
225
- user aborts do not masquerade as success. Goal/list/loop cancellation clears
226
- its timer and durable state.
223
+ - **Successful-turn reset**: a real non-error agent end on the preferred
224
+ primary clears the recovery cycle. In the current failback policy, a
225
+ successful fallback turn instead arms the durable preferred-primary probe;
226
+ `mainModelFailback=sticky` retains the historical immediate reset. Manual
227
+ model selection cancels it; host restore selections do not, and user aborts
228
+ do not masquerade as success. Goal/list/loop cancellation clears its timer
229
+ and durable state.
227
230
 
228
231
  ## Addendum v0.34.48–v0.34.56 (lifecycle/recovery hardening — the stale-handle era)
229
232
 
@@ -337,6 +340,15 @@ replacement without delivering a successor `session_start`:
337
340
  after every hour starts. Main-model recovery and detached-auditor recovery
338
341
  use this same reason-agnostic rule, with existing context/user-abort and
339
342
  safety-horizon exceptions.
343
+ - **Preferred-primary failback is durable and supervised**:
344
+ `mainModelFailback=auto` (the default) does not treat a successful fallback
345
+ turn as proof that the original primary is healthy. The recovery record keeps
346
+ `primary`, records `primaryProbeAt`, and uses
347
+ `mainModelPrimaryProbeMinutes` (15 by default) to select the primary for one
348
+ real supervised probe. A primary success clears the episode; a provider
349
+ failure walks back to the serving fallback and schedules the next reverse
350
+ probe. `sticky` preserves the legacy permanent fallback choice. The
351
+ `primaryProbeInFlight` marker and pending switch survive a session boundary.
340
352
  - **Legacy state is inert**: old `quota-waiting` phases, quota-named retry
341
353
  counters, and provider-hint fields are accepted only long enough to load
342
354
  and normalize old files. Canonical persisted state uses `retry-waiting`,
@@ -539,9 +551,24 @@ This protects against model-generated summaries losing fidelity.
539
551
  | LOW | Telegram push | v0.3.0 |
540
552
  | LOW | Sub-task auto-close | v0.3.0 |
541
553
 
554
+ ## Addendum v0.35.6 (Unattended autonomy, audit cadence, and parallelization)
555
+
556
+ - **Two-phase decision architecture (upfront grilling → zero pauses during execution)**:
557
+ - Drafting upfront is the sole interview boundary: the agent asks sharp questions about architecture, scope, error conditions, and test commands.
558
+ - Active execution is 100% unattended: the agent picks sensible architectural defaults, records rationale, and continues without interrupting the user for obvious choices or secondary questions. Non-blocking notes are deferred to the completion summary.
559
+ - Premium engineering standards: mandatory root-cause fixes, full TypeScript type safety, and comprehensive test coverage. If an approach fails verification after 2 attempts, the agent autonomously steps back and pivots to an alternative architecture.
560
+ - **Audit cadence across modes**:
561
+ - `/goal`: Evaluated by the detached isolated auditor at goal completion (`complete_goal`).
562
+ - `/list`: Evaluated by the detached isolated auditor at the completion of **every individual list task** before unlocking and activating item $N+1$. This prevents list drift, ensuring that errors in early tasks do not cascade into downstream tasks.
563
+ - `/loop`: Shell metric command evaluated on every iteration; LLM auditor does not run on intermediate iterations.
564
+ - **Parallelization architecture & opportunities**:
565
+ - *Current*: Parallel subagent exploration fan-out (spawning multiple `Explore` agents in one turn) and parallel disjoint-worktree implementation (`general-purpose` workers with `isolation: "worktree"`). Detached auditor runs concurrently in background OS process.
566
+ - *Opportunities*: Concurrent `/list` dispatch across non-conflicting tasks using isolated worktrees to eliminate head-of-line blocking for large queues, plus background contract rehearsals.
567
+
542
568
  ## Files
543
569
 
544
570
  - `docs/DESIGN.md` — **this file**
545
571
  - `README.md` — quickstart
546
572
  - `audit/pi-name-v3-registry-based.md` — naming rationale
547
573
  - `audit/pi-goal-loop-design.md` — earlier design (now superseded)
574
+
package/docs/INDEX.md CHANGED
@@ -2,11 +2,34 @@
2
2
 
3
3
  Ordered by reading path, not alphabetically.
4
4
 
5
+ ## Active focus (recent work, durable artifacts)
6
+
7
+ This package's policy contracts and recent changes are recorded in
8
+ the audit/ directory of the **repository checkout** — it is not
9
+ shipped in the npm tarball (see "Repository-only material" below).
10
+
11
+ For shipped docs, the relevant entry points are:
12
+
13
+ - `../CHANGELOG.md` — user-facing changelog; the top of the file is the
14
+ current package version. v0.35.5 adopted the six-label completion
15
+ recap; v0.35.6 added typed-boundary regression pins; v0.35.7 added
16
+ deterministic fast-fail pre-audits, zero-pause autonomous execution, and
17
+ task milestone gating; v0.35.8 added main-model preferred-primary
18
+ failback; v0.35.9 hardened cross-version npm tarball checks; v0.35.10
19
+ handles multi-entry npm dry-run reports; v0.35.11 accepts both npm report
20
+ shapes; v0.35.12 supports npm 12's keyed pack reports; v0.35.13 fixes stale-API recovery loops.
21
+ - `../README.md` — what the plugin is, install, quickstart, and the
22
+ architectural guarantee (drafting + confirm + detached auditor).
23
+ - `../INSTALL.md` — manual install / symlink setup; the recommended
24
+ companion plugins and the `auditor reads / writes are path-checked`
25
+ note.
26
+
5
27
  ## Entry points
6
28
  - `../README.md` — what the plugin is, install, quickstart
7
29
  - `../INSTALL.md` — manual install / symlink setup
8
- - `../PLAN.md` — project plan
9
- - `../LIST-PHILOSOPHY.md` — list-queue design philosophy
30
+ - `../CHANGELOG.md` — user-facing changelog; current package version is
31
+ at the top of the file (use `/glla version` to compare with the
32
+ registry).
10
33
 
11
34
  ## Architecture
12
35
  - `DESIGN.md` — plugin design (types, state, extension lifecycle)
@@ -21,11 +44,13 @@ Ordered by reading path, not alphabetically.
21
44
  - `../prompts/` — goal/loop drafting prompt templates
22
45
  - `../schemas/` — goal state JSON schema
23
46
  - `../examples/` — example objective files
24
- - `../audit/INDEX.md` — the audit trail (every shipped change, newest first)
25
47
  - `../CHANGELOG.md` — user-facing changelog (unreleased at top)
26
48
 
49
+ ## Repository-only material
50
+ The audit history and competitor research live in `audit/` and `.research/`
51
+ for contributors, but are intentionally not included in the npm tarball.
52
+
27
53
  ## Research material
28
- - `.research/` — competitor plugin sources pulled from npm tarballs for
29
- study (gitignored, local only). Re-pull with:
30
- `cd .research && npm pack <pkg> && tar xzf <tgz>` — see the
31
- positioning doc's appendix for the package list.
54
+ `.research/` — competitor plugin sources pulled from npm tarballs for study
55
+ (gitignored, local only). Re-pull with `cd .research && npm pack <pkg> &&
56
+ tar xzf <tgz>`; see the positioning doc's appendix for the package list.