pi-goal-list-loop-audit 0.37.3 → 0.38.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,29 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.38.1 — drafter/auditor fallback parity with Main (2026-09-02)
4
+
5
+ ### Fixed
6
+ Drafter model selection now emits the same `model_fallback_select` ledger as Main (`scope: drafter`) for every forbidden/unregistered skip, so `forbiddenModels` edits mid-interview are honored even when the user edits settings while the interview is open.
7
+
8
+ Drafter recovery after a provider failure now walks the dedicated 0-10 chain via `ModelSelector` with the bounded ladder (`5s → base*2^(n-1) cap 5h`, base `mainModelRetryMinutes` default 15) and `forbiddenModels` re-check, instead of an immediate unbounded loop over the lease candidates. Same-model pin (drafter pinned to the session model) remains a same-model retry and is preserved.
9
+
10
+ Auditor `model_fallback_select` ledger (`scope: auditor`) now fires for both the settings-time `resolveAuditorModel` walk and the detached audit fallback (`runAuditorFallbackWithPolicy` via `runDetachedCompletionWithFallback`), closing the gap where detached retries skipped forbidden/unregistered refs silently. `forbiddenRefs` and `retryBaseMinutes` were already wired; `onSelection` is now forwarded.
11
+
12
+ ### Changed
13
+ `Drafter fallback agents` menu row → `Drafter fallback models (up to 10)` with count prefix `N/10 · 1. ref → 2. ref` and the shared `ordered and deselectable: current drafter → fallback 1 → fallback 2…` description, matching Main and Auditor rows. Prompt title and notification strings updated to `models` vocabulary. Watchdog budgets (`auditorToolTimeoutMs 5m`, `auditorStallMs 10m` x2 cap 4x) remain independent of the ladder and unchanged.
14
+
15
+ ## 0.38.0 — event-driven supervision and drafting discipline (2026-09-02)
16
+
17
+ ### Fixed
18
+ Keep-checking supervision: scheduling is now event-driven for every plane (including `isMonitorGoal` daemons) — lifecycle/durable/child-progress signals via `ContinuousSupervisor` plus 250ms→15s adaptive fallback poll, not a guessed task-duration wait. The 120s `GLLA_MONITOR_INTERVAL_MS` throttle that made a 10s task wait up to 120s is deprecated (`void MONITOR_CHECK_INTERVAL_MS`) and kept only for display parity (`👁 MONITORING` badge). A 10s task is now picked up after ~10s even when guessed at 10m.
19
+
20
+ Pinned context-growth fixtures (`tests/context-growth-measurement.test.ts`, `tests/context-checkpoint.test.ts`) refreshed for the larger continuation prompt (23015 chars, 23123 bytes) so `npm run test:all` stays green.
21
+
22
+ ### Changed
23
+ Drafting-batches zero mid-execution questions: `buildSeedGrillMessage` now batches 2-4 sharp, seed-specific questions up front in one `ask_user_question` picker with recommended defaults per question. `LONG_RUNNING_JUDGMENT_POLICY` and `ACTIVE_EXECUTION_QUESTION_GUIDANCE` plus `prompts/goal-loop-continuation.md` enforce drafting as the only place to gather scope/acceptance — active execution targets zero further clarification unless irreversible/destructive, missing permission, or comparable-cost, picking the safest contract-preserving default otherwise and deferring preferences to the completion summary.
24
+
25
+ `docs/DESIGN-long-running-supervision.md` now documents both Now requirements and the display-only monitor contract.
26
+
3
27
  ## 0.37.3 — pi 0.84 loop stall fix — agent_start fallback for continuation start proof (2026-09-02)
4
28
 
5
29
  ### Fixed
@@ -39,6 +39,27 @@ The shared checker covers all GLLA-owned work planes: ordinary goals, list
39
39
  items and their queue, metric/spec/audit loops, detached completion auditors,
40
40
  tracked subagents, provider recovery, and lifecycle/session transitions.
41
41
 
42
+ Two Now requirements are enforced alongside the checker:
43
+
44
+ - **Keep checking, not waiting.** The checker is the primary reason to inspect
45
+ work; a guessed task duration is never a wait. A 10 s task is picked up
46
+ within the 250 ms → 15 s adaptive fallback, not after a 10 m estimate.
47
+ `isMonitorGoal` (daemon / old-goal name or age > 1 h) remains shared for
48
+ display parity (the dim 👁 MONITORING badge), but scheduling is event-driven
49
+ for every plane — the legacy 120 s monitor throttle (`GLLA_MONITOR_INTERVAL_MS`)
50
+ is retired and kept only as deprecated env-var compatibility (the checker
51
+ already backs off to 15 s when idle).
52
+ - **Zero mid-execution questions — compensate up front.** Drafting batches 2–4
53
+ sharp scope/acceptance questions with recommended defaults via a single
54
+ `ask_user_question` invocation, so active execution needs no further
55
+ clarification. During `active` execution the target is zero questions unless
56
+ proceeding would cross an irreversible/destructive external boundary, require
57
+ a missing permission/credential, or face two genuinely comparable options
58
+ with materially different outcomes. `LONG_RUNNING_JUDGMENT_POLICY` encodes the
59
+ durable-fix default; `ACTIVE_EXECUTION_QUESTION_GUIDANCE` encodes the
60
+ drafting-only discipline; `buildSeedGrillMessage` enforces the floor
61
+ (propose is blocked until the user has replied).
62
+
42
63
  ## Aggressive automation
43
64
 
44
65
  Aggressive mode is the default effective keep-going policy unless the user
package/docs/INDEX.md CHANGED
@@ -18,7 +18,7 @@ For shipped docs, the relevant entry points are:
18
18
  failback; v0.35.9 hardened cross-version npm tarball checks; v0.35.10
19
19
  handles multi-entry npm dry-run reports; v0.35.11 accepts both npm report
20
20
  shapes; v0.35.12 supports npm 12's keyed pack reports; v0.35.13 fixes stale-API recovery loops.
21
- v0.35.14–v0.37.3 continue through the supervisor freeze (`/glla pause`),
21
+ v0.35.14–v0.38.1 continue through the supervisor freeze (`/glla pause`),
22
22
  load hold, auditor picker parity, Windows launch fix, zombie-watchdog
23
23
  subagent carve-out, due-wait backstop, the `/glla agents` visibility panel,
24
24
  durable state-root selection, blank-until-resume auditor context, frozen
@@ -7,7 +7,7 @@
7
7
 
8
8
  import type { ExtensionContext } from "@earendil-works/pi-coding-agent";
9
9
 
10
- import { isForbiddenModel } from "./goal-loop-core.js";
10
+ import { appendLedger, isForbiddenModel } from "./goal-loop-core.js";
11
11
  import { MAX_MAIN_MODEL_FALLBACKS, modelRef, normalizeMainModelFallbackRefs } from "./main-model-recovery.js";
12
12
  import { ModelSelector } from "./model-selector.js";
13
13
  import type { Settings } from "./goal-settings.js";
@@ -63,15 +63,22 @@ export function resolveDrafterModel(ctx: ExtensionContext, settings: Pick<Settin
63
63
  getChain: () => configuredRefs,
64
64
  resolve: (ref) => resolveDrafterModelRef(ctx, ref),
65
65
  isForbidden: forbidden,
66
+ record: (event) => {
67
+ appendLedger(ctx.cwd, "model_fallback_select", {
68
+ scope: event.scope.kind,
69
+ fromRef: event.fromRef,
70
+ toRef: event.toRef,
71
+ reason: event.reason,
72
+ });
73
+ },
66
74
  });
67
75
  const attempted: string[] = [];
68
76
  const candidates: DrafterModelCandidate[] = [];
69
77
 
70
- // ModelSelector deliberately skips the current model so the main recovery
71
- // walker never selects the model that just failed. Drafting is different:
72
- // its configured primary may intentionally be the current session model,
73
- // and that primary still needs a lease so its later fallbacks remain
74
- // available if the first drafting turn fails.
78
+ // KEEP: drafting intentionally leases the current model when pinned (unlike
79
+ // Main which skips `current` via nextUntriedModelRef because it just failed).
80
+ // Align ledger to Main but keep this lease gate — removing it would force a
81
+ // needless setModel when the user pinned the drafter to the session model.
75
82
  const configuredPrimary = configuredRefs[0];
76
83
  if (
77
84
  configuredPrimary &&
@@ -65,11 +65,18 @@ import {
65
65
  import { BACKOFF_IDLE_RETRY_MS, HEARTBEAT_MAX_NUDGES } from "./goal-loop-backoff.js";
66
66
  import { LENGTH_CONTINUE_MAX, LENGTH_CONTINUE_TEXT } from "./length-continue.js";
67
67
 
68
+ // v0.38.0: GLLA_MONITOR_INTERVAL_MS throttling is deprecated — scheduling is
69
+ // event-driven (250ms→15s adaptive fallback) for every plane, including
70
+ // monitoring goals (see scheduleContinuation comment). Kept for env-var
71
+ // compatibility; display parity still uses isMonitorGoal, but the checker no
72
+ // longer waits a fixed 120s. The constants remain so an explicit env var does
73
+ // not silently disappear from process state during a rolling upgrade.
68
74
  const DEFAULT_MONITOR_CHECK_INTERVAL_MS = 120_000;
69
75
  const configuredMonitorIntervalMs = Number(process.env.GLLA_MONITOR_INTERVAL_MS);
70
76
  const MONITOR_CHECK_INTERVAL_MS = Number.isFinite(configuredMonitorIntervalMs) && configuredMonitorIntervalMs > 0
71
77
  ? Math.max(1_000, configuredMonitorIntervalMs)
72
78
  : DEFAULT_MONITOR_CHECK_INTERVAL_MS;
79
+ void MONITOR_CHECK_INTERVAL_MS; // deprecated throttle — scheduling is now event-driven
73
80
  import { VISION_ASSIST_GUIDANCE } from "./vision-assist.js";
74
81
  import { loadSettings } from "./goal-settings.js";
75
82
  import { clearLoopTimer, isLoopActive } from "./goal-loop.js";
@@ -1024,10 +1031,16 @@ export function scheduleContinuation(ctx: ExtensionContext, force = false, delay
1024
1031
  } catch {
1025
1032
  return;
1026
1033
  }
1027
- // v0.37.x: monitor goals (daemon, long-running >1h) check less frequently to avoid constant QUEUED churn.
1028
- if (delayMs === undefined && state.goal && isMonitorGoal(state.goal)) {
1029
- delay = Math.max(delay, MONITOR_CHECK_INTERVAL_MS);
1030
- }
1034
+ // v0.38.0 (note.md Now — "keep checking instead of waiting"): monitoring
1035
+ // goals remain visually distinct (👁 MONITORING badge via isMonitorGoal, shared
1036
+ // with the TUI), but scheduling is event-driven for every plane — the 120s
1037
+ // throttle used to delay implicit continuations for daemon/old goals and
1038
+ // made a 10s task wait up to 120s. The durable-state / lifecycle event +
1039
+ // 250ms→15s adaptive fallback in ContinuousSupervisor is the primary checker;
1040
+ // implicit continuation delay is 0 when idle, 50ms otherwise, never a guessed
1041
+ // task-duration wait. isMonitorGoal stays pure for display parity, not for
1042
+ // throttling the checker — a monitoring goal that actually finishes or
1043
+ // progresses is picked up within the fallback window, not after a fixed age.
1031
1044
  // v0.34.104 ([Image-#1]): the post-list-completion settle window delays
1032
1045
  // the first continuation after a queue auto-advance. Any real agent
1033
1046
  // activity during the window clears `postCompletionSettleUntil`, so a
@@ -783,7 +783,7 @@ export function goalArgsNeedDrafting(args: string): boolean {
783
783
  * via draftProposalBlock: propose is blocked until the user has replied.
784
784
  */
785
785
  export function buildSeedGrillMessage(tmpl: string, seed: string, tool: string): string {
786
- return `${tmpl}\n\n${LONG_RUNNING_JUDGMENT_POLICY}\n\nThe user's initial objective (verbatim): ${seed}\n\nGRILL THEM ABOUT THIS SEED BEFORE PROPOSING. ${tool} is BLOCKED until the user has replied to at least one of your questions — proposing without interviewing returns an error.\n\nHow to grill:\n- Ask ONE sharp, seed-specific question at a time — about THIS objective, not generic filler. If an ask_user_question tool is available in this session, prefer it (structured options render better); plain conversation is fine for free-form answers.\n- Every question ships with a recommended default the user can accept with "yes".\n- Probe what matters: what "done" concretely looks like (checkable evidence — files, commands, behaviors), scope boundaries (what is explicitly OUT), constraints (what must not change), and priorities when the seed bundles several wishes.\n- A non-answer ("not sure", "none", "whatever") is a trigger to offer 2-3 concrete options to pick from — never silently proceed on a non-answer.\n- Do targeted read-only research first when it makes your questions sharper (repo layout, existing docs).\n- Do NOT activate the raw seed. Do NOT implement anything. When the contract is concrete, call ${tool}.`;
786
+ return `${tmpl}\n\n${LONG_RUNNING_JUDGMENT_POLICY}\n\nThe user's initial objective (verbatim): ${seed}\n\nGRILL THEM ABOUT THIS SEED BEFORE PROPOSING. ${tool} is BLOCKED until the user has replied to at least one of your questions — proposing without interviewing returns an error.\n\nHow to grill:\n- Ask 2-4 sharp, seed-specific questions UP FRONT in ONE batched ask_user_question call when multiple unknowns exist — about THIS objective, not generic filler. Each question ships with a recommended default the user can accept with "yes" (one picker, 2-4 concrete options per question). If only one unknown remains, one focused question is fine. Prefer the structured ask_user_question picker; plain conversation is fine for free-form answers.\n- Probe what matters in that single upfront batch: what "done" concretely looks like (checkable evidence — files, commands, behaviors), scope boundaries (what is explicitly OUT), constraints (what must not change), and priorities when the seed bundles several wishes. One well-batched interview up front eliminates mid-execution interruptions — do not dribble questions out one by one during execution.\n- A non-answer ("not sure", "none", "whatever") is a trigger to offer 2-3 concrete options to pick from — never silently proceed on a non-answer.\n- Do targeted read-only research first when it makes your questions sharper (repo layout, existing docs).\n- Do NOT activate the raw seed. Do NOT implement anything. When the contract is concrete, call ${tool}.`;
787
787
  }
788
788
 
789
789
  /**
@@ -2775,7 +2775,7 @@ ${formatDurableDeferPolicyLine()}
2775
2775
  - Use an opportunistic workaround only when the durable fix is genuinely unsafe, impossible, or blocked right now; the workaround must be reversible and testable, and its durable follow-up is recorded (ledger or comment) instead of silently treated as final.
2776
2776
  - Premium engineering standards are mandatory: code must be cleanly typed, tested, architecturally sound, and resilient across lifecycle boundaries. Never lower test standards, fake assertions, or bypass types.
2777
2777
  - Autonomous pivot strategy: if an implementation approach fails verification after 2 attempts, do not loop on the same failing line. Autonomously step back, diagnose the root invariant, and pivot to a clean alternative architecture.
2778
- - Non-interruption & sensible defaults: never pause a multi-hour run for obvious choices, cosmetic naming, or non-blocking secondary questions. Pick the sensible architectural default, implement it, record the rationale, and continue. Defer non-blocking notes to the final completion summary.
2778
+ - Non-interruption & sensible defaults: never pause a multi-hour run for obvious choices, cosmetic naming, or non-blocking secondary questions. Pick the sensible architectural default, implement it, record the rationale, and continue. Defer non-blocking notes to the final completion summary. Compensate for zero mid-run questions by asking MORE up front: during drafting, batch 2-4 critical scope/acceptance questions with recommended defaults via a single ask_user_question invocation, so active execution needs no further clarification.
2779
2779
  - Decide autonomously through local implementation choices without interrupting the user. Ask one focused question ONLY at a genuine trade-off where the user's preference materially changes the outcome: an irreversible/destructive external action, a missing permission/credential, or two options with comparable real cost.
2780
2780
  - In unattended mode, choose the safest contract-preserving path and continue. If no safe choice exists, raise a concrete DECIDE question with a recommended default; never ask a vague progress question or wait on a guessed provider/quota reset.`;
2781
2781
 
@@ -2785,12 +2785,11 @@ ${formatDurableDeferPolicyLine()}
2785
2785
  * reversible local choices or turn them into user-facing pauses.
2786
2786
  */
2787
2787
  export const ACTIVE_EXECUTION_QUESTION_GUIDANCE = `ACTIVE-EXECUTION QUESTION DISCIPLINE:
2788
- - Drafting is the default place to gather scope, acceptance criteria, constraints, and trade-offs. Once active, treat the confirmed objective and verification contract as sufficient context.
2789
- - During active execution, do not ask about reversible implementation choices, naming, formatting, test shape, or whether to continue. Choose the maintainable contract-preserving option, record the rationale, and proceed.
2788
+ - Drafting is the ONLY place to gather scope, acceptance criteria, constraints, and trade-offs — batch 2-4 sharp questions up front with recommended defaults via one ask_user_question call. Once active, treat the confirmed objective and verification contract as sufficient context and do NOT reopen reversible local choices.
2789
+ - During active execution, do NOT ask about reversible implementation choices, naming, formatting, test shape, or whether to continue. Choose the maintainable contract-preserving option, record the rationale, and proceed. The target is zero mid-execution questions unless proceeding would cross an irreversible/destructive external boundary, require a missing permission/credential, or face two genuinely comparable options with materially different results.
2790
2790
  - Defer non-blocking preferences and alternatives to the completion summary (or a durable note); do not turn them into a pause or question.
2791
- - Ask one focused user question only when proceeding would cross an irreversible or destructive external boundary, requires a missing permission or credential, or two genuinely comparable options would materially change the result or acceptance.
2792
- - For a necessary question, state the exact impact, include a recommended default, and pause only the dependent action; continue independent work when possible.
2793
- - Never ask a vague progress or "what next?" question, and never wait on a guessed provider or quota reset; use bounded recovery or choose the safe default.`;
2791
+ - For the rare necessary mid-run question, state the exact impact, include a recommended default, and pause only the dependent action; continue independent work when possible.
2792
+ - Never ask a vague progress or "what next?" question, and never wait on a guessed provider or quota reset; use bounded recovery or choose the safe default. If drafting left an ambiguity, pick the safest contract-preserving default and record it rather than interrupting a multi-hour run.`;
2794
2793
 
2795
2794
  /**
2796
2795
  * v0.23.5: normalize a drafter-supplied verification contract for the
@@ -2023,8 +2023,8 @@ export function registerGoalRuntime(pi: ExtensionAPI): void {
2023
2023
  // an interview must retry that chain before the main-goal recovery path;
2024
2024
  // otherwise a drafting failure would silently consume main-model backups.
2025
2025
  if (lastA?.stopReason === "error" && draftingTarget !== null) {
2026
- const retryDrafting = (globalThis as any).handleDrafterModelFailure as ((context: ExtensionContext) => Promise<boolean>) | undefined;
2027
- if (retryDrafting) await retryDrafting(ctx);
2026
+ const retryDrafting = (globalThis as any).handleDrafterModelFailure as ((context: ExtensionContext, error?: string) => Promise<boolean>) | undefined;
2027
+ if (retryDrafting) await retryDrafting(ctx, (rawLastA as any)?.errorMessage ?? (rawLastA as any)?.error ?? lastA.text);
2028
2028
  // Whether a configured fallback was available or not, drafting owns
2029
2029
  // this error. Leave the interview open for an explicit user retry and
2030
2030
  // never pass the same failure into main-goal recovery.
@@ -904,6 +904,7 @@ export async function runDetachedCompletionWithFallback(
904
904
  onFallback?: (from: AuditorModelCandidate, to: AuditorModelCandidate, error: string) => void;
905
905
  forbiddenRefs?: readonly string[];
906
906
  retryBaseMinutes?: number;
907
+ onSelection?: (event: { scope: { kind: string }; fromRef?: string; toRef?: string; reason: string }) => void;
907
908
  } = {},
908
909
  ): Promise<{ result: DetachedAuditResult; retriedOnce: boolean; fallbackUsed: boolean; via: string }> {
909
910
  return runAuditorFallbackWithPolicy(candidates, run, {
@@ -911,6 +912,7 @@ export async function runDetachedCompletionWithFallback(
911
912
  shouldRetry: opts.shouldRetry,
912
913
  sleep: opts.sleep,
913
914
  retryBaseMinutes: opts.retryBaseMinutes,
915
+ onSelection: opts.onSelection,
914
916
  resumeCandidateRef: opts.resumeCandidateRef,
915
917
  attemptedRefs: opts.attemptedRefs,
916
918
  retryCandidateRef: opts.retryCandidateRef,
@@ -1101,6 +1103,7 @@ async function retryStoredCompletionAudit(origin: CompletionAuditOrigin = "provi
1101
1103
  shouldRetry: () => detachedAuditContext(generation, goalId, claim.attemptId!) !== null,
1102
1104
  forbiddenRefs: settings.forbiddenModels,
1103
1105
  retryBaseMinutes: settings.mainModelRetryMinutes,
1106
+ onSelection: (event: { scope: { kind: string }; fromRef?: string; toRef?: string; reason: string }) => appendLedger(liveCtx.cwd, "model_fallback_select", { scope: "auditor", fromRef: event.fromRef, toRef: event.toRef, reason: event.reason }),
1104
1107
  resumeCandidateRef: persistedAuditorCandidateRef,
1105
1108
  attemptedRefs: persistedAuditorAttemptedRefs,
1106
1109
  retryCandidateRef: persistedAuditorRetryCandidateRef,
@@ -371,6 +371,7 @@ import {
371
371
  resolveDrafterModel,
372
372
  type DrafterModelCandidate,
373
373
  } from "../drafter-model.js";
374
+ import { ModelSelector } from "../model-selector.js";
374
375
  import {
375
376
  assessSuspiciousObjective,
376
377
  deriveObjectiveFromContract,
@@ -553,25 +554,57 @@ async function restoreDrafterModel(): Promise<void> {
553
554
  }
554
555
 
555
556
  /**
556
- * Generic drafting recovery: walk the dedicated chain after any provider
557
- * error and replay a continuation of the existing interview. No error text
558
- * is classified and the main-model recovery chain is never consulted.
557
+ * Generic drafting recovery: walk the dedicated chain via ModelSelector so
558
+ * forbidden/unregistered refs are skipped with the same ledger as Main,
559
+ * and provider retries use the bounded ladder (5s → base*2^(n-1) cap 5h).
560
+ * The session model remains the last resort; same-model lease is preserved.
559
561
  */
560
- async function handleDrafterModelFailure(ctx: ExtensionContext): Promise<boolean> {
562
+ async function handleDrafterModelFailure(ctx: ExtensionContext, error?: string): Promise<boolean> {
561
563
  const lease = draftingModelLease;
562
564
  if (!lease || draftingTarget === null || lease.generation !== sessionGeneration || !extensionApi) return false;
563
- for (const candidate of lease.candidates) {
564
- if (lease.attempted.some((ref) => ref.toLowerCase() === candidate.ref.toLowerCase())) continue;
565
- lease.attempted.push(candidate.ref);
565
+ const settings = loadSettings(ctx.cwd);
566
+ const failure = classifyMainModelFailure(error);
567
+ if (failure.kind === "non-recoverable") return false;
568
+ const configuredRefs = lease.candidates.filter((candidate) => candidate.via === "configured").map((candidate) => candidate.ref);
569
+ const selector = new ModelSelector({
570
+ getChain: () => configuredRefs,
571
+ resolve: (ref) => lease.candidates.find((candidate) => candidate.ref.toLowerCase() === ref.toLowerCase())?.model,
572
+ isForbidden: (ref) => isForbiddenModel(ref, settings.forbiddenModels),
573
+ record: (event) =>
574
+ appendLedger(ctx.cwd, "model_fallback_select", {
575
+ scope: "drafter",
576
+ fromRef: event.fromRef,
577
+ toRef: event.toRef,
578
+ reason: event.reason,
579
+ }),
580
+ });
581
+ let attempt = lease.attempted.length;
582
+ for (;;) {
583
+ const pick = selector.selectNextValid({ kind: "drafter" }, lease.activeRef, lease.attempted);
584
+ for (const visited of selector.lastVisitedRefs) {
585
+ if (!lease.attempted.some((ref) => ref.toLowerCase() === visited.toLowerCase())) lease.attempted.push(visited);
586
+ }
587
+ if (!("model" in pick)) break;
588
+ const delayMs = mainModelFailureDelayMs(failure, ++attempt, (settings as any).mainModelRetryMinutes ?? 15);
589
+ if (delayMs > 5_000) await new Promise<void>((resolve) => setTimeout(resolve, delayMs));
566
590
  let switched = false;
567
- try { switched = await extensionApi.setModel(candidate.model); } catch { switched = false; }
568
- const alreadyActive = candidate.ref.toLowerCase() === lease.activeRef.toLowerCase();
591
+ try {
592
+ switched = await extensionApi.setModel(pick.model);
593
+ } catch {
594
+ switched = false;
595
+ }
596
+ const alreadyActive = pick.ref.toLowerCase() === lease.activeRef.toLowerCase();
569
597
  if (!switched && !alreadyActive) continue;
570
598
  const previous = lease.activeRef;
571
- lease.activeRef = candidate.ref;
599
+ lease.activeRef = pick.ref;
572
600
  lease.activeThinkingLevel = applyDrafterThinkingLevel(ctx, lease.requestedThinkingLevel);
573
- appendLedger(ctx.cwd, alreadyActive ? "drafter_model_retry" : "drafter_model_fallback", { from: previous, to: candidate.ref, attempted: lease.attempted });
574
- ctx.ui.notify(`Drafting provider failed; retrying the existing interview on ${candidate.ref}.`, "warning");
601
+ appendLedger(ctx.cwd, alreadyActive ? "drafter_model_retry" : "drafter_model_fallback", {
602
+ from: previous,
603
+ to: pick.ref,
604
+ attempted: lease.attempted.slice(),
605
+ delayMs,
606
+ });
607
+ ctx.ui.notify(`Drafting provider failed; retrying the existing interview on ${pick.ref} after ${Math.round(delayMs / 1000)}s.`, "warning");
575
608
  const safeSteer = (globalThis as any).safeSteerUser as ((context: ExtensionContext, text: string) => boolean) | undefined;
576
609
  if (!safeSteer || !safeSteer(ctx, DRAFTER_RECOVERY_PROMPT)) return false;
577
610
  draftingSeedInFlight = true;
@@ -554,9 +554,13 @@ export function resolveAuditorModel(
554
554
  resolve: (candidate) => tryRef(candidate).model,
555
555
  isForbidden: forbidden,
556
556
  record: (event) => {
557
+ appendLedger(ctx.cwd, "model_fallback_select", {
558
+ scope: "auditor",
559
+ fromRef: event.fromRef,
560
+ toRef: event.toRef,
561
+ reason: event.reason,
562
+ });
557
563
  if (event.reason === "forbidden" && event.toRef) {
558
- // Forbidden is an explicit user gate, not an unavailable-model
559
- // warning. Keep the forensic ledger entry but do not nudge the UI.
560
564
  appendLedger(ctx.cwd, "auditor_model_fallback", { configured: event.toRef, reason: "forbidden" });
561
565
  return;
562
566
  }
@@ -1222,13 +1226,13 @@ export async function handleSettingChoice(id: string, ctx: ExtensionContext): Pr
1222
1226
  const current = normalizeMainModelFallbackRefs(global.drafterModelFallbacks);
1223
1227
  const refs = await promptModelRefs(
1224
1228
  ctx,
1225
- `Drafter fallback agents — ordered, up to ${MAX_MAIN_MODEL_FALLBACKS}; the session agent is the last resort; forbidden models are hidden`,
1229
+ `Drafter fallback models (up to ${MAX_MAIN_MODEL_FALLBACKS}) — ordered; the session model is the last resort; forbidden models are hidden`,
1226
1230
  current,
1227
1231
  { excludeRefs: normalizeModelRefs(loadSettings(ctx.cwd).forbiddenModels), maxSelections: MAX_MAIN_MODEL_FALLBACKS, currentRef: modelRef(ctx.model) },
1228
1232
  );
1229
1233
  if (refs === undefined) return;
1230
1234
  saveSettings("global", ctx.cwd, { drafterModelFallbacks: refs.length ? refs : undefined });
1231
- ctx.ui.notify(refs.length ? `Drafter fallback agents saved: ${refs.join(" → ")}.` : "Drafter fallback agents cleared — the session model remains the last resort.", "info");
1235
+ ctx.ui.notify(refs.length ? `Drafter fallback models saved: ${refs.join(" → ")}.` : "Drafter fallback models cleared — the session model remains the last resort.", "info");
1232
1236
  return;
1233
1237
  }
1234
1238
  case "forbiddenModels": {
@@ -889,6 +889,7 @@ function registerAgentTools(pi: any): void {
889
889
  shouldRetry: () => detachedAuditContext(auditGeneration, auditGoalId, auditAttemptId) !== null,
890
890
  forbiddenRefs: settings.forbiddenModels,
891
891
  retryBaseMinutes: settings.mainModelRetryMinutes,
892
+ onSelection: (event: { scope: { kind: string }; fromRef?: string; toRef?: string; reason: string }) => appendLedger(ctx.cwd, "model_fallback_select", { scope: "auditor", fromRef: event.fromRef, toRef: event.toRef, reason: event.reason }),
892
893
  resumeCandidateRef: persistedAuditorCandidateRef,
893
894
  attemptedRefs: persistedAuditorAttemptedRefs,
894
895
  retryCandidateRef: persistedAuditorRetryCandidateRef,
@@ -322,12 +322,12 @@ export function buildSettingsRows(
322
322
  {
323
323
  id: "drafterModelFallbacks",
324
324
  section: "drafter",
325
- label: "Drafter fallback agents",
325
+ label: `Drafter fallback models (up to ${MAX_MAIN_MODEL_FALLBACKS})`,
326
326
  valueText: settings.drafterModelFallbacks?.length
327
- ? modelChainText(settings.drafterModelFallbacks, drafterThinking, subagent)
328
- : "none (session last resort)",
327
+ ? `${settings.drafterModelFallbacks.length}/${MAX_MAIN_MODEL_FALLBACKS} · ${settings.drafterModelFallbacks.map((ref, index) => `${index + 1}. ${modelThinkingText(ref, drafterThinking, subagent)}`).join(" → ")}`
328
+ : `0/${MAX_MAIN_MODEL_FALLBACKS} · ${sessionRef} · ${drafterThinking} (last resort)`,
329
329
  sourceText: src("drafterModelFallbacks"),
330
- description: "ordered drafting-only fallback agents; each shows its effective/requested thinking level when the model registry exposes capabilities",
330
+ description: "ordered and deselectable: current drafter → fallback 1 → fallback 2…; every recoverable provider failure switches one eligible fallback at a time",
331
331
  },
332
332
  );
333
333
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-goal-list-loop-audit",
3
- "version": "0.37.3",
3
+ "version": "0.38.1",
4
4
  "description": "Mission control for autonomous pi: interview-drafted goals, an audited task queue, and forever-loops (metric, spec, project-audit) that run for hours. A detached extension-less auditor process re-verifies every completion with raw evidence without holding the main pi turn; confirmed drafts, decision pauses and consent gates keep you in charge.",
5
5
  "license": "AGPL-3.0-only",
6
6
  "author": "dracon",
@@ -73,7 +73,7 @@ When the agent calls any of these, the orchestrator tracks the call and persists
73
73
  - **Auditor rehearsal**: When the verification contract has checks a subagent can re-run, spawn ONE fresh-context `reviewer` agent to rehearse the contract before calling `complete_goal`.
74
74
  - **Eager continuation.** When in doubt, KEEP GOING on sub-tasks. If a subagent fails, retry with a different approach. Don't ask permission to continue — just continue. Pause only when you are genuinely blocked on information that does not exist in the repo, or the user explicitly pauses you.
75
75
  - **Premium engineering & autonomous pivot strategy.** Always implement root-cause architectural fixes rather than superficial band-aids or test hacks. If an implementation approach fails tests after 2 attempts, do NOT loop on the same failing line: autonomously step back, diagnose the root invariant, and pivot to an alternative clean architecture.
76
- - **Non-interruption & sensible defaults law.** Upfront drafting is where you interview the user; once the goal is active, you are in UNATTENDED autonomous mode. Never pause a multi-hour goal for obvious decisions, naming preferences, or non-blocking secondary questions. Choose the sensible architectural default, implement it, record the rationale, and continue. Defer non-blocking notes to the final completion summary.
76
+ - **Non-interruption & sensible defaults law.** Batch 2–4 sharp questions UP FRONT in drafting (one `ask_user_question` picker with recommended defaults per question — scope, done-criteria, constraints, priorities) so active execution needs zero further clarification. Once the goal is ACTIVE, you are in UNATTENDED autonomous mode: never pause a multi-hour goal for obvious decisions, naming preferences, or non-blocking secondary questions. Compensate for zero mid-run questions by asking more upfront. Choose the sensible architectural default, implement it, record the rationale, and continue. Defer non-blocking notes to the final completion summary.
77
77
  - **Bound every long command.** Wrap test suites, builds, and dev servers in `timeout <seconds>` (e.g. `timeout 120 bun test src/lib`). An unbounded command that hangs burns an hour; a bounded one burns two minutes and tells you it hung. If a command produces no output for many minutes, treat it as hung: kill it, diagnose why, rerun bounded.
78
78
  - **Chunk output near context-full & microcompaction.** When the conversation is heavy (long-running audit, deep debug, big rollout), prefer smaller commits, smaller tool outputs, and focused reasoning — one or two punchy paragraphs, one well-scoped tool call at a time. Don't try to fit a thousand lines of work into one reply. Spool massive stdout/diffs to disk logs if needed. glla's auto-continue fires on `stop_reason="length"` and will reschedule you; chunking is cheaper than recovering from the cap. Save large file writes for their own turns.
79
79