pi-goal-list-loop-audit 0.38.71 → 0.38.91
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +101 -0
- package/README.md +7 -3
- package/docs/ARCHITECTURE.md +116 -0
- package/docs/DESIGN.md +112 -0
- package/docs/INDEX.md +3 -1
- package/docs/PROMOTION-CONTRACT.md +67 -0
- package/docs/SETTINGS.md +2 -1
- package/extensions/goal-commands.ts +193 -5
- package/extensions/goal-loop-auditor-process.ts +186 -9
- package/extensions/goal-loop-auditor.ts +1 -1
- package/extensions/goal-loop-core.ts +49 -4
- package/extensions/goal-loop-display.ts +15 -4
- package/extensions/goal-loop-shield.ts +4 -1
- package/extensions/goal-loop-stats.ts +237 -2
- package/extensions/goal-loop-subagents.ts +0 -3
- package/extensions/goal-settings.ts +20 -0
- package/extensions/loops/goal-activation.ts +134 -65
- package/extensions/loops/goal-auditor-hooks.ts +19 -4
- package/extensions/loops/goal-orchestrator.ts +0 -1
- package/extensions/loops/goal-runtime-globals.ts +0 -93
- package/extensions/loops/goal-session.ts +0 -18
- package/extensions/loops/goal-settings-ui.ts +23 -4
- package/extensions/loops/goal-tools.ts +92 -5
- package/extensions/loops/goal-ui.ts +2 -14
- package/extensions/loops/goal.ts +37 -1
- package/extensions/quota-retry.ts +0 -2
- package/extensions/settings-menu.ts +9 -0
- package/package.json +5 -2
- package/schemas/goal.schema.json +16 -0
- package/scripts/goal-auditor-worker.mjs +222 -32
- package/scripts/release-pack-smoke.mjs +6 -2
- package/scripts/run-tests.d.mts +6 -0
- package/scripts/run-tests.mjs +71 -0
- package/skills/glla-delegate/SKILL.md +7 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,106 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.38.91 — Tool descriptions enumerate plan/audit/verify/add (2026-09-21)
|
|
4
|
+
|
|
5
|
+
- `/goal plan`, `/goal audit`, `/goal verify`, `/list plan`, and `/list add` existed but were undiscoverable from the tool descriptions. Both descriptions now enumerate them with one-line roles; a pin keeps the enumerations (audit/POST-090-IMPROVEMENT-SWEEP-2026-09-21.md).
|
|
6
|
+
|
|
7
|
+
## 0.38.90 — Monitoring visuals: standby reasons go dim (2026-09-21)
|
|
8
|
+
|
|
9
|
+
- Paused-on-standby cards rendered background-agent status narration in warning yellow like errors. Standby reasons now render dim like waits and cap at 2 wrapped rows (audit/MONITORING-VISUALS-2026-09-21.md).
|
|
10
|
+
|
|
11
|
+
## 0.38.89 — UI polish finale: stats usage, header prefix, status→timeline pointer (2026-09-21)
|
|
12
|
+
|
|
13
|
+
- `/glla stats <typo>` now names the usage (`outcomes | challenges | premature | json | project=<path>`) instead of silently rendering the default table; the JSON path stays total for machine readers.
|
|
14
|
+
- The stats table header carries the `/glla` prefix like the stats error paths (`/glla stats …`).
|
|
15
|
+
- The `/goal status` card ends with a `Trail: /goal timeline` pointer to its slow twin.
|
|
16
|
+
|
|
17
|
+
## 0.38.88 — Pi-shim isolation: per-file process-state reset (2026-09-21)
|
|
18
|
+
|
|
19
|
+
- New `__testOnlyResetProcessState()` composite invokes all 17 latch resets, and the test preload calls it before each file — per-file isolation is now structural instead of 99 files hand-picking resets (21 picked none). Membership is pinned: a future reset that bypasses the composite fails the suite (audit/PI-SHIM-ISOLATION-2026-09-21.md).
|
|
20
|
+
|
|
21
|
+
## 0.38.87 — Globals retirement slice 2: 24 ambient slots deleted (2026-09-21)
|
|
22
|
+
|
|
23
|
+
- The runtime-globals registry drops 207 → 183: 11 `__testOnly*` hooks, `classifySessionHandleInvalidation`, `ownerFilePath`, `recentActions`, `COMPACT_FIRST_NUDGE_PERCENT`, three settings-UI helpers, and six pin-only constants — every consumer already imports directly, so only the ambient slots delete. Two test files switch from the globalThis path to direct imports (audit/GLOBALS-RETIREMENT-SLICE-2-2026-09-21.md).
|
|
24
|
+
|
|
25
|
+
## 0.38.86 — Docs: architecture overview + promotion contract (2026-09-21)
|
|
26
|
+
|
|
27
|
+
- New `docs/ARCHITECTURE.md`: GLLA in 20 minutes — entry/runtime, the three loops, the audit lifecycle, persistence, and the test suite, every claim grounded in source (stage names, handler counts, and ledger-type counts verified, not remembered).
|
|
28
|
+
- New `docs/PROMOTION-CONTRACT.md`: the list item → goal → archive seam diagrammed and pinned — activate, archive, auto-advance, abort/no-advance, and the one-active-thing rule.
|
|
29
|
+
- Repair: v0.38.85's `timeline` subcommand broke the no-hardcoded-guidance pin (the `/goal` tool description is an allowed surface map; its allowlist didn't know the new subcommand). The full-suite gate caught it; the pin now tracks the enumeration.
|
|
30
|
+
|
|
31
|
+
## 0.38.85 — `/goal timeline`: one what-happened + next-action view (2026-09-21)
|
|
32
|
+
|
|
33
|
+
- New `/goal timeline [N]` renders the goal's own trail — goal-scoped ledger events merged with auditHistory verdicts in time order, key types humanized (pauses, tiers, shields, consents), unknown types compact but never hidden — plus a footer naming the single next action from live state (decision to answer, resume time, open objections, or working). Read-only like `/glla log` (audit/TIMELINE-2026-09-21.md).
|
|
34
|
+
|
|
35
|
+
## 0.38.84 — Mode check learns /loop: recurring seeds route to loops (2026-09-21)
|
|
36
|
+
|
|
37
|
+
- `crossRecommendMode` now catches open-ended/recurring seeds (cadence words, monitor/watch, keep-X-under-Y) in goal and list drafting and steers the interview toward `/loop` — neither goals (end on approval) nor list items (close once) fit "keep doing it". Bounded-until phrasing ("until done/green") stays a goal, and aggregate seeds still win (audit/LOOP-INFERENCE-2026-09-21.md).
|
|
38
|
+
|
|
39
|
+
## 0.38.83 — Lifecycle map: session_start stages extracted + numbered (2026-09-21)
|
|
40
|
+
|
|
41
|
+
- The 800-line `session_start` callback is now navigable: three verbatim stages extracted into named nested helpers (`admitSessionStart`, `claimSessionRootOrNotify`, `retentionSweepAuditJobs`) and all 12 stages carry numbered banners. Behavior-identical — moved code is byte-identical except the early returns (audit/LIFECYCLE-MAP-2026-09-21.md).
|
|
42
|
+
- Premise correction: the ×3 loop wiring this item assumed does not exist — one entry, one `registerGoalRuntime`, single registrations; the per-handler gates differ deliberately and the structure is intentionally source-pinned. Full extraction would fight load-bearing pins, so the map documents the stages instead of moving them.
|
|
43
|
+
|
|
44
|
+
## 0.38.82 — Fast/slow suite split: `npm test` runs the fast set (2026-09-21)
|
|
45
|
+
|
|
46
|
+
- The full serialized suite takes ~7 minutes — slow enough to skip. `npm test` now runs the fast set via `scripts/run-tests.mjs`: everything minus the 12 slowest files (evidence-timed in `tests/slow-files.mjs`, behavioral-orchestrator alone is ~90s). `npm run test:slow` runs those files, `npm run test:changed` runs git-affected files, and `npm run test:all` / `release:check` still run the whole suite — nothing escapes the gate (audit/TEST-SPLIT-2026-09-21.md).
|
|
47
|
+
- Explicit paths always win: naming a slow file runs it (bun ignore patterns would otherwise beat explicit selection). `-t` filters pass through without tripping path detection.
|
|
48
|
+
- Repair: `package-lock.json` had drifted to 0.38.78 behind three version bumps; re-synced (the `glla-version` sync test caught it on the first fast run).
|
|
49
|
+
|
|
50
|
+
## 0.38.81 — Risk-tiered auditing: light tier, spot-checks, full-audit consent (2026-09-21)
|
|
51
|
+
|
|
52
|
+
- Every audit now dispatches at a risk tier. Full tier = audit plus the falsification round (today's behavior); light tier = the same single-round audit with the same brief, shield, and tool floor, only round 2 skipped. Escalation-only: rework history, activity ceilings, high-stakes language, draft-time `fullAudit` consent, and `complete_goal requestFullAudit` can each push a claim UP to full; nothing pushes it down (audit/RISK-TIERS-2026-09-21.md).
|
|
53
|
+
- Spot-checks: a sampled fraction of light audits (setting `auditSpotCheckRate`, default 0.1, 0 = off) silently runs full so under-tiering is caught statistically. Tier + spot mark record on every verdict; `/glla stats challenges` gains the light count and spot flip rate for calibration.
|
|
54
|
+
- Invariants: light never means none, and the agent can demand full but never light. Tier decisions ledger as `audit_tier_decided` with reasons.
|
|
55
|
+
|
|
56
|
+
## 0.38.80 — Challenge outcomes recorded; `/glla stats challenges` (2026-09-21)
|
|
57
|
+
|
|
58
|
+
- The worker's falsification outcome (`confirmed` / `flipped` / `not-applicable` / `skipped:<reason>`) now threads through the parent boundary onto every recorded `AuditVerdict` (both settle sites), so audit quality becomes measurable instead of believed (audit/AUDIT-METRICS-2026-09-21.md).
|
|
59
|
+
- New `/glla stats challenges` view (composes with `json`/`project=`): challenged/confirmed/flipped/skipped counts plus the flip rate over challenged runs only. Skipped challenges and legacy verdicts stay unknown, never zero-filled.
|
|
60
|
+
|
|
61
|
+
## 0.38.79 — Schema covers runToDone; T6 green (2026-09-20)
|
|
62
|
+
|
|
63
|
+
- `runToDone` joins `schemas/goal.schema.json`: v0.38.73 added the consent flag to the `Goal` interface but not to the published schema, and the T6 drift test caught it in the full suite (2457 pass / 1 fail). One additive boolean property, zero behavior change; full serial suite re-run green (audit/SCHEMA-RUNTO-DONE-2026-09-20.md).
|
|
64
|
+
|
|
65
|
+
## 0.38.78 — Remove 6 dead exports; repair hourly-pin window (2026-09-20)
|
|
66
|
+
|
|
67
|
+
- Free-only scope pass: 6 exported consts with zero references anywhere (including their own files) removed — `DEFAULT_HOURLY_RETRY_PROBE`, 3 `EXPLORE_DEFAULT_*` aliases, `MAX_AUTOMATIC_QUOTA_RETRY_SEC`, `isSubagentQuotaResult`. A first sweep wrongly dropped 49 own-file-used consts; caught by review before any release, fully restored, redone with the correct predicate (audit/FREE-SCOPE-PASS-2026-09-20.md).
|
|
68
|
+
- Hourly source-pin window 28k → 30k: the v0.38.73 run-to-done consent block pushed both schedule calls to ~28.2k/28.5k. Calls unchanged and correctly placed; the pin's own comments document this maintenance.
|
|
69
|
+
|
|
70
|
+
## 0.38.77 — Retire 15 single-module runtime globals (222 → 207, 2026-09-20)
|
|
71
|
+
|
|
72
|
+
- First slice of the `globalThis` bridge retirement: 15 registry names used in exactly one module and nowhere else (no cross-module, test, script, or top-level consumers) lose their registration, ambient declaration, and typed entry. Owning modules keep their locals untouched — zero runtime behavior change (audit/GLOBALS-RETIREMENT-SLICE-2026-09-20.md).
|
|
73
|
+
- Causality note: the `glla-status-ux` + `behavioral-orchestrator` combo shows 2 pre-existing order-dependent failures (v0.35.15 pause/resume notification counts); verified identical on the pre-retirement tree via worktree comparison.
|
|
74
|
+
|
|
75
|
+
## 0.38.76 — Auditor challenge round: approvals earn a falsification pass (2026-09-20)
|
|
76
|
+
|
|
77
|
+
- The detached worker now runs a second bounded RPC round in a fresh session whenever round 1 approves, with an adversarial brief (re-verify load-bearing claims; disapprove with specifics on any genuine gap). Outputs compose by final line: a challenge disapproval flips the verdict, a re-confirm preserves it. Disapprovals, impossibles, and failures stay single-round (audit/AUDITOR-CHALLENGE-2026-09-20.md).
|
|
78
|
+
- Fail-open, recorded: a failed challenge truncates back to byte-identical round-1 output and records `skipped:<reason>` in `result.challenge` (`confirmed`/`flipped`/`not-applicable` otherwise). Cancellation still wins outright. New `challenging` progress phase with HUD labels.
|
|
79
|
+
- Tradeoff: worst-case approval latency roughly doubles (same bounds per round); parent wall-timeout + retry machinery absorbs overruns as before.
|
|
80
|
+
|
|
81
|
+
## 0.38.75 — Compaction survival suite; cost ceiling verified pre-existing (2026-09-20)
|
|
82
|
+
|
|
83
|
+
- New `tests/compaction-survival.test.ts`: three genuine mid-run `session_compact` events followed by full detached-audit completion, plus a compact around an in-flight audit. Previously only projection shape, in-flight suppression, and settle probes were covered — nothing proved a goal still completes after real compacts (audit/COMPACTION-SURVIVAL-2026-09-20.md).
|
|
84
|
+
- Cost-ceiling review: no build needed — per-goal token limits already pause with notice (opt-in via limit > 0) and loop token budgets already stop. Verified by inspection, not changed.
|
|
85
|
+
|
|
86
|
+
## 0.38.74 — `/glla stats outcomes`: completion metrics per project (2026-09-20)
|
|
87
|
+
|
|
88
|
+
- New `outcomes` view for `/glla stats` (composes with `json`/`project=`): done/aborted/open counts, completion rate, mean audit rounds to approval, mean wall-clock hours and tokens per completed goal, plus the run-to-done vs supervised completion split — so run-to-done mode can be judged on numbers, not belief (audit/STATS-OUTCOMES-2026-09-20.md).
|
|
89
|
+
- Unknowns stay unknown: goals without timing/usage data are skipped from means, never zero-filled; pre-flag goals land in the supervised-or-legacy bucket.
|
|
90
|
+
|
|
91
|
+
## 0.38.73 — Run-to-done mode: draft up front, carry to completion (2026-09-20)
|
|
92
|
+
|
|
93
|
+
- New per-goal run-to-done consent at draft time: the interview asks supervised vs run-to-done, the Confirm dialog discloses the mode (auto session resume + immediate decision auto-default until complete or a hard stop), and consent is durable on the goal plus ledgered (`run_to_done_consented`). Single-goal drafts only in v1; batch/list drafts refuse the flag with guidance (audit/RUN-TO-DONE-2026-09-20.md).
|
|
94
|
+
- Runtime: run-to-done goals auto-resume held work at session start without global autoResume (per-goal consent wins), auto-default every decision immediately (`run_to_done_auto_default`, same autoDefaultLog mechanics), and treat audit caps as hard stops (park, no TODO conversion). Blocked-on-external pauses still park; the auditor is never skipped; user abort always wins. Status line carries a `run to done` chip.
|
|
95
|
+
- `docs/INDEX.md` version trail fixed (was pinned at v0.38.71).
|
|
96
|
+
|
|
97
|
+
## 0.38.72 — Automatic audit-job retention sweep at session start (2026-09-20)
|
|
98
|
+
|
|
99
|
+
- The proven-dead audit-job sweep now runs automatically at admitted-owner session start (previously manual `/glla audits health cleanup` only, which let 207 job dirs accumulate). Windowed by the existing `auditJobRetentionMs` setting, ledgered as `audit_jobs_retention_sweep` when it reaps, fail-silent by design (audit/AUDIT-JOB-RETENTION-SWEEP-2026-09-20.md).
|
|
100
|
+
- Pre-convention finished audits (result.json on file, no worker lock ever written) now classify dead past the retention window instead of ambiguous forever; present-but-corrupt locks and unfinished lockless dirs stay ambiguous for operator inspection.
|
|
101
|
+
- `paused-suspicious-close` settlement ceiling 10s to 30s: the close path legitimately takes ~9.4s unloaded, so the old ceiling flaked under any load.
|
|
102
|
+
- Stale `goal-loop-shield.ts` header corrected (runner section documented as shell-free, matching helpers named as the pure subset).
|
|
103
|
+
|
|
3
104
|
## 0.38.71 — Decision cards lead with the action, plaque verbs deduped (2026-09-20)
|
|
4
105
|
|
|
5
106
|
- Decision pause cards render the saved/action row before the numbered options, so Pi core tail truncation cuts late options and history instead of the resume path (UI survey finding 1, audit/UI-SURVEY-2026-09-20.md).
|
package/README.md
CHANGED
|
@@ -485,9 +485,13 @@ npm run check
|
|
|
485
485
|
npm run release:check
|
|
486
486
|
```
|
|
487
487
|
|
|
488
|
-
`npm
|
|
489
|
-
|
|
490
|
-
|
|
488
|
+
`npm test` runs the fast set (the full serialized suite minus the 12
|
|
489
|
+
slowest files — ~4 minutes instead of ~7). `npm run test:slow` runs
|
|
490
|
+
those slow files, `npm run test:changed` runs only git-affected files,
|
|
491
|
+
and `npm run test:all` runs everything. `npm run release:check` runs
|
|
492
|
+
the serialized Bun suite, TypeScript, the jiti reproduction, offline
|
|
493
|
+
auditor-extension validation, and npm pack. The test count changes as
|
|
494
|
+
regressions are added; the useful result is `0 fail`.
|
|
491
495
|
|
|
492
496
|
For design rationale, see [`docs/DESIGN.md`](docs/DESIGN.md). For the shipped
|
|
493
497
|
document index, see [`docs/INDEX.md`](docs/INDEX.md). For publishing, see
|
|
@@ -0,0 +1,116 @@
|
|
|
1
|
+
# Architecture: GLLA in 20 minutes
|
|
2
|
+
|
|
3
|
+
Read top to bottom. Each section names the files that own it; the
|
|
4
|
+
[promotion contract](PROMOTION-CONTRACT.md) covers the list→goal→archive
|
|
5
|
+
seam in full.
|
|
6
|
+
|
|
7
|
+
## 0. The one paragraph (1 min)
|
|
8
|
+
|
|
9
|
+
GLLA is mission control for long-running Pi work: it drafts objectives
|
|
10
|
+
with the user, runs them as audited goals, and verifies every completion
|
|
11
|
+
with a detached auditor that never shares the worker's session. Three
|
|
12
|
+
loops (`/goal`, `/list`, `/loop`) share one runtime, one ledger
|
|
13
|
+
(`.pi-glla/active.jsonl`), and one rule: **one active thing at a time**.
|
|
14
|
+
|
|
15
|
+
## 1. Entry and runtime (3 min)
|
|
16
|
+
|
|
17
|
+
- Entry: `extensions/loops/goal.ts` → `registerGoalRuntime`
|
|
18
|
+
(`extensions/loops/goal-activation.ts`). All 17 Pi event
|
|
19
|
+
registrations (`session_start`, `agent_end`, `tool_result`, …) live
|
|
20
|
+
there; per-handler gates differ deliberately (worker vs subagent
|
|
21
|
+
vs foreign vs host-successor planes).
|
|
22
|
+
- `session_start` is the lifecycle spine: 12 numbered stages from
|
|
23
|
+
admission → root ownership → retention sweep → rebind reset →
|
|
24
|
+
queue convergence → recovery prep → resume-consent → load barrier
|
|
25
|
+
→ arbitration → held-loop resume → continuation release →
|
|
26
|
+
stored-claim release. Three closed stages are extracted helpers;
|
|
27
|
+
the rest stay inline and source-pinned on purpose (see
|
|
28
|
+
`audit/LIFECYCLE-MAP-*.md`).
|
|
29
|
+
- Shared state: `extensions/goal-state.ts` (`state`, `replaceState`)
|
|
30
|
+
plus module-level runtime vars; cross-file sharing goes through
|
|
31
|
+
`goal-runtime-globals.ts` (being retired slice by slice).
|
|
32
|
+
|
|
33
|
+
```
|
|
34
|
+
Pi events ──▶ registerGoalRuntime ──▶ state + ledger ──▶ UI (status line, banners, /goal status)
|
|
35
|
+
│
|
|
36
|
+
▼
|
|
37
|
+
detached auditor (separate process)
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
## 2. The three loops (5 min)
|
|
41
|
+
|
|
42
|
+
- **Loop 1 `/goal`** — one objective, drafted (interview + Confirm) or
|
|
43
|
+
started, worked until the agent claims completion, then audited.
|
|
44
|
+
States: `active → auditing → complete|aborted`, with `paused`
|
|
45
|
+
overlays (decision, wait, error, blocked, standby).
|
|
46
|
+
- **Loop 2 `/list`** — a durable queue of short items. Items activate
|
|
47
|
+
head-first into real goals (`policy: "list"`), complete → archive →
|
|
48
|
+
auto-advance; aborts archive without advancing; nothing ever
|
|
49
|
+
returns to the queue. Full contract: PROMOTION-CONTRACT.md.
|
|
50
|
+
- **Loop 3 `/loop`** — metric loops (`LoopState` in
|
|
51
|
+
`goal-loop-forever.ts`: target, measure command, direction,
|
|
52
|
+
iteration/max/plateau/stall accounting). No `Goal`, no audit
|
|
53
|
+
verdicts — it ends on bounds, plateau, or `/loop stop`. Metricless
|
|
54
|
+
spec loops end on bounds only.
|
|
55
|
+
|
|
56
|
+
One-active-thing: a live loop blocks list activation (loudly, with
|
|
57
|
+
the way out); completion cascades hand the surface to the next item.
|
|
58
|
+
|
|
59
|
+
## 3. The audit lifecycle (5 min)
|
|
60
|
+
|
|
61
|
+
```
|
|
62
|
+
complete_goal → claim → dispatch (tier) → worker (round 1 [+ challenge]) → shield → verdict → history → archive|park
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
- **Claim**: `complete_goal` persists a `PendingCompletion` (summaries,
|
|
66
|
+
Left-out, finding groups, gates, optional full-audit request).
|
|
67
|
+
- **Dispatch** (`goal-loop-auditor-process.ts`): resolves the risk
|
|
68
|
+
tier (full = audit + falsification round; light = single round),
|
|
69
|
+
ledgers `audit_tier_decided`, spawns the worker with a hashed
|
|
70
|
+
`request.json`.
|
|
71
|
+
- **Worker** (`scripts/goal-auditor-worker.mjs`): runs Pi `--no-session`
|
|
72
|
+
against the brief, parses the final-line verdict (`<approved/>` /
|
|
73
|
+
`<disapproved/>`), challenges approvals in a fresh session unless
|
|
74
|
+
the tier says light. Fail-open: a failed challenge yields
|
|
75
|
+
byte-identical round-1 output, recorded as `skipped:*`.
|
|
76
|
+
- **Parent**: validates attempt/hash/revision, enforces the audit-tool
|
|
77
|
+
floor and the regression shield (contract items must be cited),
|
|
78
|
+
records the `AuditVerdict` (tier, challenge outcome, duration),
|
|
79
|
+
then archives on clean approval or parks with objections on
|
|
80
|
+
disapproval. Spot-checks (default 1-in-10 light audits) calibrate
|
|
81
|
+
the tiers; `/glla stats challenges` reports the flip rates.
|
|
82
|
+
|
|
83
|
+
## 4. Persistence (3 min)
|
|
84
|
+
|
|
85
|
+
- **Ledger** (`.pi-glla/active.jsonl`): append-only JSONL, ~340 event
|
|
86
|
+
types, the forensic trail. Stats, timeline, and recovery all read it.
|
|
87
|
+
- **State**: snapshots persist the live goal/loop/list; `readState`
|
|
88
|
+
reconciles with the ledger at boundaries.
|
|
89
|
+
- **Sidecars**: durable per-queue-item files; deleted before an item
|
|
90
|
+
leaves the queue so half-moves can't resurrect work.
|
|
91
|
+
- **Archive** (`.pi-glla/archive/<id>.md`): exclusive-create,
|
|
92
|
+
intent-journaled, human + machine record. Terminal and immutable.
|
|
93
|
+
- **Recovery**: crash-safe by construction — intent journals,
|
|
94
|
+
archive fences, revision-bound verdicts, stale-refusal instead of
|
|
95
|
+
silent overwrite. When in doubt it parks loudly, never proceeds
|
|
96
|
+
quietly.
|
|
97
|
+
|
|
98
|
+
## 5. The test suite (3 min)
|
|
99
|
+
|
|
100
|
+
- ~260 files, must run **serialized** (`--parallel=1
|
|
101
|
+
--max-concurrency=1`): parallel files trip Bun's nesting guard.
|
|
102
|
+
- `npm test` = fast set (minus the 12 slowest, see
|
|
103
|
+
`tests/slow-files.mjs`); `npm run test:slow`, `test:changed`,
|
|
104
|
+
`test:all` cover the rest. `release:check` = full suite + tsc +
|
|
105
|
+
jiti repro + offline auditor-extension check + pack + smoke.
|
|
106
|
+
- Two test kinds: **behavioral** (MockPi harness drives real code)
|
|
107
|
+
and **source pins** (`assert.match` on runtime source — order and
|
|
108
|
+
presence guards for load-bearing structure). Pins are curated; if
|
|
109
|
+
a refactor breaks one, the pin's *intent* decides whether the
|
|
110
|
+
refactor or the pin is wrong.
|
|
111
|
+
|
|
112
|
+
## Where to go next
|
|
113
|
+
|
|
114
|
+
- Operating: `README.md`, `docs/SETTINGS.md`, `/goal timeline`.
|
|
115
|
+
- Deep design: `docs/DESIGN.md` (+ addenda), `docs/RELEASING.md`.
|
|
116
|
+
- History of why: `audit/*.md` — every incident and decision, dated.
|
package/docs/DESIGN.md
CHANGED
|
@@ -673,6 +673,118 @@ shapes (details in CHANGELOG.md; each is pinned by tests):
|
|
|
673
673
|
|
|
674
674
|
- **Primary scope is the cwd project**: `listAuditCollectTarget`, `projectAuditTarget`, and `auditTarget` now state "current project rooted at the cwd where pi was opened (treat any nested .git as a separate project boundary — do not walk into parent or sibling projects)". The TIGHT scout brief is "named directories under cwd" — external code outside cwd may be READ only to diagnose a failure that blocks the current project, and a finding about external code is valid only when it affects the current project (a typo in an unrelated sibling project is out of scope and never auto-queued). This closes the "audit the parent when you opened a subproject" leak observed when hellhunter was audited from the dracon-platform root and vice-versa.
|
|
675
675
|
|
|
676
|
+
## Addendum v0.38.72 (retention policy)
|
|
677
|
+
|
|
678
|
+
- **audit-jobs**: time-window retention, not count-capped. `auditJobRetentionMs`
|
|
679
|
+
(default 15m, max 7d) bounds how long finished audit transcripts stay
|
|
680
|
+
readable; the admitted owner sweeps proven-dead and pre-convention finished
|
|
681
|
+
dirs past the window at every session start (ledgered when it reaps).
|
|
682
|
+
A count cap was rejected: it could delete recent evidence, while the time
|
|
683
|
+
window already bounds growth (worst case ≈ one week of audits).
|
|
684
|
+
- **ledger segments**: append-only history, no rotation. Rotation moves the
|
|
685
|
+
live ledger into segments; segments themselves are never compacted or
|
|
686
|
+
deleted. Disk cost is the audit trail's price and stays visible in
|
|
687
|
+
`.pi-glla/ledger-segments/`.
|
|
688
|
+
|
|
689
|
+
## Addendum v0.38.73 (run-to-done mode)
|
|
690
|
+
|
|
691
|
+
- **Draft up front, carry to the end.** A goal draft may carry run-to-done
|
|
692
|
+
consent: the interview asks supervised vs run-to-done, the Confirm
|
|
693
|
+
dialog discloses the mode in full view (the Confirm IS the consent),
|
|
694
|
+
and the flag is durable on the goal plus ledgered (`run_to_done_consented`).
|
|
695
|
+
- **What the consent grants**: session-start auto-resume of held work
|
|
696
|
+
without global autoResume (specific beats general), and immediate
|
|
697
|
+
decision auto-default (`run_to_done_auto_default`, same autoDefaultLog
|
|
698
|
+
mechanics as the budget path). Waits already auto-continue; quota still
|
|
699
|
+
sleeps to reset.
|
|
700
|
+
- **Hard stops still park**: audit disapproval caps (no TODO conversion —
|
|
701
|
+
repeated rejection means the approach is wrong), consecutive-error
|
|
702
|
+
ceilings, provider outage, user abort/pause. Blocked-on-external pauses
|
|
703
|
+
also still park: missing user input is a hard-stop-class dependency,
|
|
704
|
+
not friction. The auditor is never skipped.
|
|
705
|
+
- **Visibility**: the status line carries a `run to done` chip — an
|
|
706
|
+
auto-running goal must never look supervised. Single-goal drafts only
|
|
707
|
+
in v1; list-queue items stay supervised.
|
|
708
|
+
|
|
709
|
+
## Addendum v0.38.74 (outcome metrics)
|
|
710
|
+
|
|
711
|
+
- **`/glla stats outcomes`** aggregates what the ledger already records:
|
|
712
|
+
done/aborted/open, completion rate, mean audit rounds to approval,
|
|
713
|
+
mean wall-clock hours and tokens per completed goal, and the
|
|
714
|
+
run-to-done vs supervised completion split. Unknowns stay unknown
|
|
715
|
+
(skipped from means, never zero-filled); pre-flag goals land in the
|
|
716
|
+
supervised-or-legacy bucket. The point is comparative: mode and
|
|
717
|
+
process changes get judged on numbers.
|
|
718
|
+
|
|
719
|
+
## Addendum v0.38.76 (auditor challenge round)
|
|
720
|
+
|
|
721
|
+
- **Approvals earn a falsification pass.** The worker runs a second bounded
|
|
722
|
+
RPC round in a fresh session (no anchoring on round-1 reasoning) with an
|
|
723
|
+
adversarial brief. Final-line verdict composition: challenge disapproval
|
|
724
|
+
flips, re-confirm preserves; the shield keeps reading round-1's evidence
|
|
725
|
+
block (first match wins). Only approvals challenge — disapprovals
|
|
726
|
+
already force rework.
|
|
727
|
+
- **Fail-open, recorded.** A failed challenge truncates to byte-identical
|
|
728
|
+
round-1 output (`result.challenge: skipped:<reason>`); cancellation
|
|
729
|
+
wins outright (`ok:false`, no fallback). Worst-case approval latency
|
|
730
|
+
roughly doubles; parent timeout/retry absorbs overruns.
|
|
731
|
+
|
|
732
|
+
## Addendum v0.38.80 (audit metrics)
|
|
733
|
+
|
|
734
|
+
- **Challenge outcomes are recorded, then judged.** The falsification
|
|
735
|
+
outcome threads onto every `AuditVerdict`, and `/glla stats
|
|
736
|
+
challenges` reports confirmed/flipped/skipped plus the flip rate over
|
|
737
|
+
challenged runs. Same unknowns-stay-unknown rule as outcomes: skipped
|
|
738
|
+
rounds and legacy verdicts never join the denominator. The flip rate
|
|
739
|
+
is the calibration signal for risk-tiered auditing: it says whether
|
|
740
|
+
round 2 earns its latency.
|
|
741
|
+
|
|
742
|
+
## Addendum v0.38.81 (risk-tiered auditing)
|
|
743
|
+
|
|
744
|
+
- **Tiers differ by one round.** Full = audit plus falsification round;
|
|
745
|
+
light = the identical single-round audit (same brief, shield, tool
|
|
746
|
+
floor), round 2 skipped via `challenge: false` in the worker request.
|
|
747
|
+
Light never means none — that invariant is the design's spine.
|
|
748
|
+
- **Escalation-only resolution.** The pure `resolveAuditTier` accumulates
|
|
749
|
+
full-tier reasons (draft consent, agent request, rework history,
|
|
750
|
+
activity ceilings, high-stakes vocabulary) and nothing pushes a claim
|
|
751
|
+
down. Agent-written text is read for escalation only. Computed per
|
|
752
|
+
dispatch through the shared `resolveClaimAuditTier` so both dispatch
|
|
753
|
+
paths tier identically; every decision ledgers as
|
|
754
|
+
`audit_tier_decided` with reasons.
|
|
755
|
+
- **Spot-checks close the loop.** `auditSpotCheckRate` (default 0.1)
|
|
756
|
+
silently escalates sampled light audits; verdicts carry the tier and
|
|
757
|
+
spot mark, and the challenges view reports the spot flip rate — the
|
|
758
|
+
calibration signal for the v1 ceilings and vocabulary.
|
|
759
|
+
|
|
760
|
+
## Addendum v0.38.82 (fast/slow suite split)
|
|
761
|
+
|
|
762
|
+
- **The everyday loop is fast; the gate is whole.** `npm test` excludes
|
|
763
|
+
the 12 evidence-timed slow files via repeated bun
|
|
764
|
+
`--path-ignore-patterns` (comma-separated does not union); `test:slow`
|
|
765
|
+
runs them; `test:all` and `release:check` still run everything.
|
|
766
|
+
Explicit test paths drop the exclusion — naming a file means running
|
|
767
|
+
it. The slow list (`tests/slow-files.mjs`) carries per-file timings
|
|
768
|
+
and is validity-checked by `tests/test-split.test.ts`.
|
|
769
|
+
|
|
770
|
+
## Addendum v0.38.83 (lifecycle map)
|
|
771
|
+
|
|
772
|
+
- **Navigate, don't restructure.** The `session_start` callback's 12
|
|
773
|
+
stages are numbered in banners; the three closed/narrow stages
|
|
774
|
+
(admission gate, root ownership, retention sweep) moved verbatim
|
|
775
|
+
into named helpers. The per-handler gates differ on purpose and the
|
|
776
|
+
structure is intentionally source-pinned (including a char-window
|
|
777
|
+
pin) — the pins are load-bearing curation against restructures
|
|
778
|
+
that regress, so further stages stay inline and mapped.
|
|
779
|
+
|
|
780
|
+
## Addendum v0.38.85 (goal timeline)
|
|
781
|
+
|
|
782
|
+
- **One screen answers "what happened, what's next".** `/goal timeline`
|
|
783
|
+
merges the goal's ledger trail with its audit verdicts in time
|
|
784
|
+
order (unscoped legacy events claimed by lifetime window, unknown
|
|
785
|
+
types compact-never-hidden) and derives the single next action
|
|
786
|
+
from live state. Read-only surface like `/glla log`.
|
|
787
|
+
|
|
676
788
|
## Files
|
|
677
789
|
|
|
678
790
|
- `docs/DESIGN.md` — **this file**
|
package/docs/INDEX.md
CHANGED
|
@@ -16,7 +16,7 @@ Policy contracts and recent changes live in the `audit/` directory of the
|
|
|
16
16
|
failback; v0.35.9 hardened cross-version npm tarball checks; v0.35.10
|
|
17
17
|
handles multi-entry npm dry-run reports; v0.35.11 accepts both npm report
|
|
18
18
|
shapes; v0.35.12 supports npm 12's keyed pack reports; v0.35.13 fixes stale-API recovery loops.
|
|
19
|
-
v0.35.14–v0.38.
|
|
19
|
+
v0.35.14–v0.38.91 continue through the supervisor freeze (`/glla pause`),
|
|
20
20
|
load hold, auditor picker parity, Windows launch fix, zombie-watchdog
|
|
21
21
|
subagent carve-out, due-wait backstop, the `/glla agents` visibility panel,
|
|
22
22
|
durable state-root selection, blank-until-resume auditor context, frozen
|
|
@@ -46,6 +46,8 @@ Policy contracts and recent changes live in the `audit/` directory of the
|
|
|
46
46
|
the registry); post-release work may appear in `Unreleased` above it.
|
|
47
47
|
|
|
48
48
|
## Architecture
|
|
49
|
+
- `ARCHITECTURE.md`: 20-minute newcomer overview (three loops, audit lifecycle, persistence)
|
|
50
|
+
- `PROMOTION-CONTRACT.md`: the list item → goal → archive seam, diagrammed
|
|
49
51
|
- `DESIGN.md`: plugin design (types, state, extension lifecycle)
|
|
50
52
|
- `DESIGN-long-running-supervision.md`: v0.36.0 event/progress-driven supervision, aggressive recovery, terminal recaps, and future decision checklist
|
|
51
53
|
- `GLLA-POSITIONING-AND-DECOMPOSITION-2026-08-08.md`: ecosystem
|
|
@@ -0,0 +1,67 @@
|
|
|
1
|
+
# The promotion contract: list item → goal → archive
|
|
2
|
+
|
|
3
|
+
The one seam every GLLA user must understand. A queue item is inert
|
|
4
|
+
text; a goal is a live audited objective; an archive is an immutable
|
|
5
|
+
record. Promotion moves exactly one way, and every edge below is
|
|
6
|
+
grounded in `activateNextListItem` (`extensions/loops/goal-list-queue.ts`)
|
|
7
|
+
and `archiveCurrentGoal` (`extensions/loops/goal-orchestrator.ts`).
|
|
8
|
+
|
|
9
|
+
```
|
|
10
|
+
QUEUED (.pi-glla queue + sidecar) LIVE (state.goal, policy="list") ARCHIVED (.pi-glla/archive/<id>.md)
|
|
11
|
+
─────────────────────────────── ────────────────────────────────── ──────────────────────────────────
|
|
12
|
+
inert: objective + contract audited: turns, claims, verdicts immutable: markdown + machine record
|
|
13
|
+
│ │ ▲
|
|
14
|
+
│ /list next · cascade · list_activate │ complete (auditor approves) │
|
|
15
|
+
│ ─────────────────────────────────────────▶ ├────────────────────────────────────────────┤
|
|
16
|
+
│ guards: loop must not own surface; │ cascade: parent group closes if last │
|
|
17
|
+
│ suspicious items get a repair card; │ subtask done, then next item activates; │
|
|
18
|
+
│ sidecar deleted BEFORE the item leaves │ empty queue → "List complete" │
|
|
19
|
+
│ │ │
|
|
20
|
+
│ activation FAILS (setGoal refused) │ abort (/goal cancel, /list next, …) │
|
|
21
|
+
│ ◀───────────────────────────────────────── ├────────────────────────────────────────────┤
|
|
22
|
+
│ item + sidecar RESTORED to the queue │ NO cascade (aborts pick their own step); │
|
|
23
|
+
│ │ item NEVER returns to the queue │
|
|
24
|
+
│ disapproval │ │
|
|
25
|
+
│ ─ ─ ─ (queue untouched) ─ ─ ─ ▶ │ rework in place; queue waits │
|
|
26
|
+
```
|
|
27
|
+
|
|
28
|
+
## The rules
|
|
29
|
+
|
|
30
|
+
1. **Activation takes the item out of the queue.** `takeAt` removes
|
|
31
|
+
it from RAM and the disk sidecar is deleted first — a half-moved
|
|
32
|
+
item cannot reappear as pending work after reload. What activates
|
|
33
|
+
is a real `Goal` with `policy: "list"`, carrying the item's
|
|
34
|
+
contract, agent role, subtask binding, and repair target.
|
|
35
|
+
2. **Guards refuse loudly, except the head-group skip.** A live loop,
|
|
36
|
+
a suspicious objective, or an undeletable sidecar refuses with a
|
|
37
|
+
banner and a ledger event. The one silent step is skipping a head
|
|
38
|
+
*group* to its first open child (groups are containers, not work)
|
|
39
|
+
— ledgered as `list_group_auto_skipped`, never bannered.
|
|
40
|
+
3. **Completion cascades; abort does not.** An approved list goal
|
|
41
|
+
archives, closes its parent group when it was the last open
|
|
42
|
+
subtask, and auto-activates the next item. An aborted goal
|
|
43
|
+
archives as aborted and stops — auto-advancing on abort once
|
|
44
|
+
double-activated (v0.2.0), so aborts pick their own next step.
|
|
45
|
+
4. **Nothing returns to the queue.** There is no requeue path: once
|
|
46
|
+
activated, the item's queue position is gone. Abort archives it
|
|
47
|
+
as aborted; to retry it, re-add it. The sole exception is a
|
|
48
|
+
*failed* activation (the prior live objective could not archive),
|
|
49
|
+
which restores item + sidecar to the queue.
|
|
50
|
+
5. **Disapproval never touches the queue.** Rework happens in place
|
|
51
|
+
on the live goal; the queue waits behind it.
|
|
52
|
+
6. **Standalone goals hand off, then stop.** A completed `/goal`
|
|
53
|
+
with a waiting list activates the head item; an aborted one
|
|
54
|
+
does not. The archive (`.pi-glla/archive/<id>.md`, exclusive
|
|
55
|
+
create, intent-journaled) is the durable record either way.
|
|
56
|
+
|
|
57
|
+
## Where to look
|
|
58
|
+
|
|
59
|
+
- Activation: `activateNextListItem` in
|
|
60
|
+
`extensions/loops/goal-list-queue.ts` (one-active-thing gate,
|
|
61
|
+
group scan, suspicious-objective repair, sidecar discipline,
|
|
62
|
+
setGoal-failure restore).
|
|
63
|
+
- Terminal: `archiveCurrentGoal` + the cascade in
|
|
64
|
+
`extensions/loops/goal-orchestrator.ts` (archive fence, intent
|
|
65
|
+
journal, parent close, advance, list-complete notice).
|
|
66
|
+
- The trail: `/goal timeline` narrates any live goal's journey
|
|
67
|
+
through these states.
|
package/docs/SETTINGS.md
CHANGED
|
@@ -30,7 +30,7 @@ copies are ignored (the recovery runtime reads the global file):
|
|
|
30
30
|
`autoResume`, `drafterModel`, `drafterThinkingLevel`,
|
|
31
31
|
`drafterModelFallbacks`, `compactorModel`, `compactorModelFallbacks`,
|
|
32
32
|
`auditorModelFallbacks`, `auditorToolTimeoutMs`, `auditorStallMs`,
|
|
33
|
-
`auditJobRetentionMs`, `auditorInspection`.
|
|
33
|
+
`auditJobRetentionMs`, `auditSpotCheckRate`, `auditorInspection`.
|
|
34
34
|
|
|
35
35
|
## Keys
|
|
36
36
|
|
|
@@ -58,6 +58,7 @@ copies are ignored (the recovery runtime reads the global file):
|
|
|
58
58
|
| `auditorToolTimeoutMs` | `300000` | Base budget per auditor tool call (30s–6h). Global-only. |
|
|
59
59
|
| `auditorStallMs` | `600000` | Base silence budget for the detached auditor (1m–24h). Global-only. |
|
|
60
60
|
| `auditJobRetentionMs` | `900000` | How long proven-dead audit job dirs are kept (0–7d, 0 = reap now). Global-only. |
|
|
61
|
+
| `auditSpotCheckRate` | `0.1` | Fraction of light-tier audits silently escalated to full (0 = off, 1 = calibrate). Global-only. |
|
|
61
62
|
| `auditorInspection` | `false` | Auditor runs as a persistent session you can tail/resume. Global-only. |
|
|
62
63
|
| `notifyCmd` | unset | Shell command on goal complete / pause / loop stop; message is `$1`. |
|
|
63
64
|
| `tokenLimit` | unset (off) | Per-goal token budget; crossing it pauses. `0` = off. |
|