docks-kit 0.17.1 → 0.17.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (87) hide show
  1. package/AGENTS.md +41 -33
  2. package/cli/docs/omp-context.md +104 -0
  3. package/cli/docs/omp-models.md +106 -56
  4. package/cli/src/argv.ts +38 -469
  5. package/cli/src/argvSurface.ts +371 -0
  6. package/cli/src/argvValidate.ts +122 -0
  7. package/cli/src/commands/docs.ts +62 -42
  8. package/cli/src/commands/harnesses.ts +38 -45
  9. package/cli/src/commands/model.ts +48 -53
  10. package/cli/src/commands/models.ts +24 -26
  11. package/cli/src/commands/omp.ts +132 -112
  12. package/cli/src/commands/plugins.ts +20 -20
  13. package/cli/src/commands/skills.ts +22 -20
  14. package/cli/src/commands/status.ts +87 -80
  15. package/cli/src/commands/sync.ts +113 -87
  16. package/cli/src/commands/toolchain.ts +21 -21
  17. package/cli/src/commands/update.ts +139 -117
  18. package/cli/src/efforts.ts +35 -34
  19. package/cli/src/engine-native/DESIGN.md +30 -30
  20. package/cli/src/engine-native/bun.ts +77 -55
  21. package/cli/src/engine-native/claudeLsp.ts +154 -0
  22. package/cli/src/engine-native/claudeOptionalPlugins.ts +113 -0
  23. package/cli/src/engine-native/claudePluginPasses.ts +348 -0
  24. package/cli/src/engine-native/claudeRemovals.ts +195 -0
  25. package/cli/src/engine-native/claudeRetired.ts +4 -4
  26. package/cli/src/engine-native/claudeRuntime.ts +94 -68
  27. package/cli/src/engine-native/claudeSettings.ts +174 -0
  28. package/cli/src/engine-native/claudeSettingsModifiers.ts +50 -54
  29. package/cli/src/engine-native/claudeSync.ts +176 -456
  30. package/cli/src/engine-native/codexConfig.ts +158 -0
  31. package/cli/src/engine-native/codexHooks.ts +91 -0
  32. package/cli/src/engine-native/codexPlugins.ts +290 -0
  33. package/cli/src/engine-native/codexStatus.ts +24 -0
  34. package/cli/src/engine-native/codexSync.ts +110 -568
  35. package/cli/src/engine-native/codexToml.ts +140 -138
  36. package/cli/src/engine-native/deps.ts +144 -107
  37. package/cli/src/engine-native/engineCtx.ts +126 -0
  38. package/cli/src/engine-native/exec.ts +98 -87
  39. package/cli/src/engine-native/failures.ts +3 -3
  40. package/cli/src/engine-native/harnesses.ts +62 -72
  41. package/cli/src/engine-native/index.ts +34 -315
  42. package/cli/src/engine-native/jq.ts +25 -24
  43. package/cli/src/engine-native/logger.ts +98 -100
  44. package/cli/src/engine-native/models.ts +50 -45
  45. package/cli/src/engine-native/modes.ts +83 -80
  46. package/cli/src/engine-native/ompFileDeploy.ts +111 -0
  47. package/cli/src/engine-native/ompMarketplace.ts +136 -0
  48. package/cli/src/engine-native/ompOverlay.ts +67 -64
  49. package/cli/src/engine-native/ompPaths.ts +49 -38
  50. package/cli/src/engine-native/ompPlugins.ts +170 -0
  51. package/cli/src/engine-native/ompRemovals.ts +159 -0
  52. package/cli/src/engine-native/ompSync.ts +69 -409
  53. package/cli/src/engine-native/ompYaml.ts +43 -41
  54. package/cli/src/engine-native/os/darwin.ts +5 -5
  55. package/cli/src/engine-native/os/index.ts +15 -15
  56. package/cli/src/engine-native/os/linux.ts +5 -5
  57. package/cli/src/engine-native/os/posix.ts +28 -24
  58. package/cli/src/engine-native/os/targets.ts +17 -17
  59. package/cli/src/engine-native/os/types.ts +31 -31
  60. package/cli/src/engine-native/os/windows.ts +61 -61
  61. package/cli/src/engine-native/parseArgs.ts +122 -372
  62. package/cli/src/engine-native/parseHelp.ts +66 -0
  63. package/cli/src/engine-native/parseModifiers.ts +222 -0
  64. package/cli/src/engine-native/services.ts +34 -34
  65. package/cli/src/engine-native/settings.ts +15 -15
  66. package/cli/src/engine-native/sharedTypes.d.ts +46 -0
  67. package/cli/src/engine-native/skillsInstall.ts +105 -0
  68. package/cli/src/engine-native/skillsLinks.ts +203 -0
  69. package/cli/src/engine-native/skillsManifest.ts +16 -0
  70. package/cli/src/engine-native/skillsPrune.ts +108 -0
  71. package/cli/src/engine-native/skillsSync.ts +28 -362
  72. package/cli/src/engine-native/syncConcurrency.ts +55 -0
  73. package/cli/src/engine-native/syncDispatch.ts +141 -0
  74. package/cli/src/engine-native/toolchain.ts +65 -62
  75. package/cli/src/engine.ts +65 -57
  76. package/cli/src/generated/sotPayload.ts +7 -7
  77. package/cli/src/kitHome.ts +44 -40
  78. package/cli/src/main.ts +63 -56
  79. package/cli/src/manifests.ts +53 -54
  80. package/cli/src/md.d.ts +2 -2
  81. package/cli/src/payload.ts +9 -9
  82. package/cli/src/services.ts +21 -13
  83. package/cli/tsconfig.json +4 -1
  84. package/docks-kit +13 -2
  85. package/docks-kit.ps1 +21 -6
  86. package/package.json +10 -2
  87. package/cli/src/engine-native/claudePlugins.ts +0 -510
package/AGENTS.md CHANGED
@@ -28,11 +28,16 @@ launcher can fall back to Bun source.
28
28
 
29
29
  | Path | Purpose |
30
30
  |------|---------|
31
- | `docks-kit` / `docks-kit.ps1` | POSIX and Windows CLI launchers. On supported hosts, each runs the matching binary in `cli/dist/` only when its `--version` matches `package.json`. Otherwise it runs Bun-from-source and auto-installs Bun plus `node_modules`. Hosts outside the support matrix fail before source fallback. The standalone platform release binary provides no-Bun recovery. |
31
+ | `docks-kit` / `docks-kit.ps1` | POSIX and Windows CLI launchers. On supported hosts, each runs the matching binary in `cli/dist/` only when its `--version` matches `package.json`. Otherwise it runs Bun-from-source and auto-installs Bun plus `node_modules`, and the mismatch warning names the rebuild remedy. A launcher never deletes a binary: on a host without Bun it is the only remaining recovery path. Hosts outside the support matrix fail before source fallback. The standalone platform release binary provides no-Bun recovery. |
32
+ | `cli/build-binaries.sh` | Compiles the six release binaries. It checksums the six known artifact names rather than globbing `docks-kit-*`, so a packed `docks-kit-<version>.tgz` is never attested as a release binary. It stamps `cli/dist/VERSION` only when the directory holds one version, and on the next build discards every artifact that stamp attributes to a different version, because a launcher refuses to run one and keeping it makes `SHA256SUMS` span two versions. An artifact with no stamp has unknown provenance and may be a hand-built or downloaded recovery binary, so the script keeps it and warns; `--prune` discards those too. Nothing here removes an artifact the stamp vouches for. |
32
33
  | `cli/src/engine-native/` | EngineNative implementation for `sync`, `model`, and `toolchain`; idempotent, flag-gated for destructive reconciliation |
33
34
  | `cli/src/engine-native/ompSync.ts` | omp file deployment, marketplace registration, and plugin synchronization |
34
35
  | `cli/src/commands/omp.ts` | `docks-kit omp` session launcher: renders the free-model run overlay and forwards args to omp |
35
36
  | `cli/src/engine-native/ompOverlay.ts` | Free-model overlay render plus catalog parse and thinking-ceiling helpers |
37
+ | `cli/src/engine-native/<axis>.ts` | One change axis per file: `codexConfig`/`codexHooks`/`codexPlugins`/`codexStatus`, `claudeSettings`/`claudeRemovals`/`claudePluginPasses`/`claudeOptionalPlugins`/`claudeLsp`, `parseModifiers`/`parseHelp`, `ompFileDeploy`/`ompMarketplace`/`ompPlugins`, `skillsManifest`/`skillsLinks`/`skillsInstall`/`skillsPrune`, `engineCtx`/`syncConcurrency`/`syncDispatch`; orchestrators keep prior exports |
38
+ | `cli/src/engine-native/sharedTypes.d.ts` | Single definition per shared shape (manifest records, settings edits, omp session model, scalar modifier flags) |
39
+ | `.oxlintrc.json` / `.oxfmtrc.json` | Mechanical lint (correctness, suspicious, eqeqeq, no-var, prefer-const, no-unused) and code format scope |
40
+ | `cli/test/fixtures/` | Golden fixtures: `home-fresh`, `home-drift`, `home-invalid-json`, `codex-toml`, `statusline` |
36
41
  | `cli/` | Effect 4 RC CLI + bundled docs topics |
37
42
  | `SoT/models.json` | Kit-verified Claude and Codex model catalog |
38
43
  | `SoT/toolchain.json` | Toolchain floors manifest (verified pins consumed by EngineNative) |
@@ -66,15 +71,18 @@ omp SoT notes:
66
71
  - `SoT/.omp/AGENTS.md`, `config.yml`, `models.yml`, and `mcp.json` deploy to `~/.omp/agent/`.
67
72
  - `SoT/.omp/intercom.json` deploys to `$PI_CODING_AGENT_DIR/intercom/config.json`. The default root is `~/.pi/agent`.
68
73
  - `ompSync.ts syncMergedYaml` deep-merges `config.yml` through `ompYaml.ts mergeOmpConfig` and `models.yml` through `mergeOmpModels`. Both wrap one generic mapping merge; only the config wrapper prunes stale `retry.fallbackChains` wildcards.
69
- - `cycleOrder` ends with the `fable` role (`modelRoles.fable` = `anthropic/claude-fable-5-1:medium`, `modelTags.fable` visible, `retry.fallbackChains.fable` empty), so the model switcher reaches Fable 5.1 as its fourth stop and never falls back off it. The hidden `switch_fable` role points at the same model and level.
70
- - `modelRoles.task` is `openai-codex/gpt-6-astra:low`, chosen for 2.60 s TTFT and 4k output tokens per task at cost parity with the previous `gpt-5.6-sol:high`. `task.agentModelOverrides` points four reviewer names at `@task`; only the bundled `reviewer` and `security-reviewer` are discoverable OMP agents, so those two inherit `task`, and the `code-reviewer` and `plan-reviewer` entries are dormant. Artificial Analysis publishes no per-level Astra Coding Agent Index score, so reviewer output is the signal to watch.
71
- - `SoT/.omp/models.yml` deploys to `~/.omp/agent/models.yml` through `mergeOmpModels`, which keeps every deployed-only key. A user file may carry provider credentials, so whole-file replacement is wrong. It restricts the `gpt-6-astra` effort ladder to `low`, so every Astra subagent runs `low` while `task.maxEffort: high` lets `scout` and `sonic` reach Luna `high` with `effort: hi`. `cli/docs/omp-models.md` records the verification.
74
+ - `ompSync.ts` runs `ompRemovals.ts syncOmpRemovals, retired-key inventory` right after the config merge, and that pass force-prunes retired kit-owned keys from `~/.omp/agent/config.yml` on every sync, without `--reconcile`. The pass is required because `mergeOmpConfig` is additive, so removing a key from the SoT alone never removes it from a deployed file. A key retired with a recorded value is pruned only while the deployed value still matches that value, so a user edit survives; a key retired outright, such as `providers.webSearchOrder`, is pruned at any value.
75
+ - `cycleOrder` ends with `astra` as its fifth stop. `modelRoles.astra` is `openai-codex/gpt-6-astra:xhigh`, and `modelTags.astra` is visible. Astra and Fable fall back to each other through concrete selectors. `modelRoles.fable` remains `anthropic/claude-fable-5-1:medium`, with visible `modelTags.fable`. The hidden `switch_fable` role uses the same Fable selector and keeps an empty fallback chain.
76
+ - `modelRoles.task` is `openai-codex/gpt-5.6-sol:high`. The four reviewer entries in `task.agentModelOverrides` inherit it through `@task`. Only bundled `reviewer` and `security-reviewer` are discoverable OMP agents; `code-reviewer` and `plan-reviewer` stay dormant.
77
+ - `modelRoles.web` is `web/firecrawl` and `retry.fallbackChains.web` carries the explicit 27-entry provider order. The kit declares both keys because the legacy `providers.webSearchOrder` key is retired: omp expands it in memory into these two keys and then drops it, and never writes that expansion back to disk. An explicit chain replaces omp's built-in web order wholesale, so the list must stay complete.
78
+ - `SoT/.omp/models.yml` declares Astra's full `low, medium, high, xhigh, max` ladder with `defaultLevel: xhigh` as the worked provider ladder-override example. `ompYaml.ts mergeOmpModels` preserves deployed-only keys in `~/.omp/agent/models.yml`, because a user file may carry provider credentials. Whole-file replacement is wrong.
72
79
  - `SoT/.omp/AGENTS.md` carries the rule `Please remove all mannered prose.` Anthropic's Fable 5.1 prompting guide documents mannered prose as a Fable 5.1 behavior and gives that sentence as its short-version fix: https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1
73
80
  - `cli/docs/omp-models.md` (topic `omp-models`) records the role map rationale and the Artificial Analysis snapshot behind it. Model choices change with published benchmarks, so update that topic in the same commit as a role change.
81
+ - `SoT/.omp/config.yml` sets `compaction.thresholdTokens` to `-1`, omp's schema default sentinel selecting reserve-based behavior: the trigger becomes `contextWindow` minus `max(floor(contextWindow * 0.15), 16384)`, which is 231,200 on the 272,000-token Codex window and 850,000 on the 1,000,000-token Anthropic window. The key ships as `-1` rather than being deleted because `ompYaml.ts mergeOmpConfig` is additive, so removing a key from SoT never removes it from a deployed file. `cli/docs/omp-context.md` (topic `omp-context`) carries the derivation and the measured evidence, so update that topic in the same commit as any compaction-setting change.
74
82
  - Sync registers the `docks` marketplace. It installs or upgrades `docks@docks` and `plan-lifecycle@docks` at user scope.
75
83
  - Sync installs `pi-intercom` at the verified version from `SoT/toolchain.json`.
76
84
  - The omp CLI is upstream-owned and self-updating through `omp update`. Sync never installs or upgrades the CLI.
77
- - `docks-kit omp [--model <selector>|--pick] [args...]` starts one interactive omp session on a single free model. It renders a run overlay to `~/.cache/docks-kit/omp-free-<model>-<digest>.yml` (mode 0600) and passes it through omp's repeatable `--config` flag. `ompOverlay.ts overlayFileName` gives each model its own file, because omp can re-read the overlay during a live session and a second launcher on another model must not rewrite it. Deployed `~/.omp/agent/` files stay untouched, so the next plain `omp` run uses the paid configuration again. Remaining arguments forward verbatim to omp.
85
+ - `docks-kit omp [--model <selector>|--pick] [args...]` starts one interactive omp session with every model role set to one free model and every retry chain empty. Higher-precedence model selection can replace those overlay values. It renders the overlay to `~/.cache/docks-kit/omp-free-<model>-<digest>.yml` (mode 0600) and passes it through omp's repeatable `--config` flag. `ompOverlay.ts overlayFileName` gives each model its own file, because omp can re-read the overlay during a live session and a second launcher on another model must not rewrite it. Deployed `~/.omp/agent/` files stay untouched, so the next plain `omp` run uses the paid configuration again. Remaining arguments forward verbatim to omp.
78
86
  - The session model persists per machine in `~/.docks-kit/state.json` under `ompSession`, next to `harnesses`. The default is `opencode-zen/muse-spark-1.3-contributor-free` at `xhigh`. `ompOverlay.ts ladderCeiling, advisorLevelFor` derive both levels from the ladder of the chosen model for the `--model` path, and `ompOverlay.ts planEffortChoice` drives the `--pick` wizard, which asks whether every role shares one level and otherwise takes one level for the main roles and one for the advisor from that model's own ladder. Levels never come from a fixed list: free ladders are not uniform, three free models publish no `medium`, and two publish no ladder, which the overlay renders as bare selectors. The picker lists only zero-cost catalog models.
79
87
 
80
88
  For per-tool SoT layouts (`SoT/.claude/`, `SoT/.codex/`, `SoT/.omp/`), see the matching SoT directory.
@@ -90,8 +98,10 @@ For per-tool SoT layouts (`SoT/.claude/`, `SoT/.codex/`, `SoT/.omp/`), see the m
90
98
  - **Per-machine harness selection.** `~/.docks-kit/state.json` drives a flag-less sync. A missing file selects `claude`, `codex`, and `agents`; it never selects `omp` implicitly. `sync` never prompts and never writes the selection file. `docks-kit harnesses` is the only command that writes the `harnesses` selection. The same state file also carries the independent `ompSession` key, which only `docks-kit omp` writes; every writer merges over the stored record, so neither key can drop the other.
91
99
  - **Additive by default.** Keys present in deployed config but absent from SoT are preserved on default sync. This protects user-only additions, but means drift accumulates — neither flag-less reset can clean it up. The one exception is the Claude `removed` manifest (`claudeSync.ts REMOVED_MANIFEST, baseline removal inventory`), a curated list of unambiguous kit-owned artifacts that `claudeSync.ts syncRemovals, baseline artifact prune` force-prunes on every sync, including the home-relative `~/.local/bin/session-relay` artifact installed outside `~/.claude`; see `CLAUDE.md` § Pruning stale artifacts.
92
100
  - **`--reconcile` / `--prune` are the kit-owned reconcile flags.** Orthogonal — `--reconcile` reconciles the settings layer (SoT-declared keys/tables/arrays win; user-only keys and nested objects are preserved; permissions arrays are replaced wholesale by SoT). `--prune` uninstalls kit-managed installations not in the SoT (plugins, marketplaces, and `~/.agents/skills/*` entries tracked in `~/.agents/.kit-managed-skills`). Combine for a full reset to SoT's kit-managed scope. User-only additions outside the kit's scope (custom env vars, mcpServers, manually-installed skills, third-party plugins not declared in SoT) are always preserved. Each tool's per-tool file documents the specific paths and diff recipes.
93
- - **SOLID-aligned modules.** `cli/src/engine-native/parseArgs.ts` owns flag parsing and validation. `toolchain.ts` owns verified-version floor reporting over `SoT/toolchain.json`; `bun.ts` owns the shared, memoized Bun bootstrap; `claudeRuntime.ts` owns Claude settings materialization. `claudeSync.ts`, `codexSync.ts`, and `skillsSync.ts` own their tool-specific sync logic. `ompSync.ts` owns omp file and plugin sync. `ompYaml.ts` owns the omp YAML merge. `harnesses.ts` owns per-machine harness selection. `index.ts` is the thin orchestrator. The public CLI seam is `cli/src/engine.ts`.
101
+ - **SOLID-aligned modules.** Axis modules own one concern each (see layout); orchestrators (`claudeSync.ts`, `codexSync.ts`, `skillsSync.ts`, `ompSync.ts`, `index.ts`, `argv.ts`, `parseArgs.ts`) keep prior exports. New shared shapes go in `sharedTypes.d.ts`, never a second local copy. The public CLI seam is `cli/src/engine.ts`.
94
102
  - **Small, reviewable changes.** Bundled multi-concern PRs are harder to review and revert. Split an engine/CLI change and a per-tool config change unless the change requires atomicity.
103
+ - **Version bumps.** Bump `package.json` on user-visible behavior change; bump a `SoT/toolchain.json` `verified` pin only after testing that release. Refactors, splits, and doc edits bump nothing.
104
+ - **Regenerate the payload.** After `SoT/`, `notification.mp3`, `package.json`, or `cli/tsconfig.json` changes, run `bun cli/scripts/generate-sot-payload.ts`; `bun run check:generated` verifies it. After file moves, clamp cited `metadata.source_files[].lines` ranges so `skillMetadata.test.ts` passes.
95
105
  - **Dry-run before destructive flags.** Always preview with `./docks-kit sync --dry-run` (or the relevant `diff <(jq -S …)` recipe in the per-tool file) before invoking `--reconcile` or `--prune`. User-added permissions / env vars / plugins absent from SoT will be discarded.
96
106
  - **SoT prompt files are rules, not explanation.** `SoT/.claude/CLAUDE.md`, `SoT/.codex/AGENTS.md`, and `SoT/.omp/AGENTS.md` are loaded into agent sessions' prompt context — every line costs prompt tokens on every turn for every user. Restrict their content to rules, heuristics, and `<constraint>` blocks the agent must *act on* during a turn. Do NOT add inline source citations (`Source: …`, attributed quotes), "why this rule exists" preface text, version-watermarking trivia (e.g. "Distilled from X v2.0, captured 2025-11-07"), per-bug workarounds, or installation instructions. Provenance, motivation, and historical context belong in `CLAUDE.md` / `AGENTS.md` at the repo root (humans read once) or in commit messages — never in the SoT. For every line, apply the official test: …
97
107
  - **Cache-invariance for kit-authored prompt surfaces.** Never put timestamps, counters, or mutable state into SoT prompt files, hook outputs that land in the cached prefix, or tool definitions — cache breaks force cold-start writes. Dynamic context belongs in runtime-injected messages (e.g. SessionStart hook output), which is exactly how the kit's date/config injection works.
@@ -109,7 +119,7 @@ For per-tool SoT layouts (`SoT/.claude/`, `SoT/.codex/`, `SoT/.omp/`), see the m
109
119
  - **One exemption: the kit's own package.** `install.sh` and `install.ps1` end with `bun add -g docks-kit@latest`, because a global installer that pinned itself would install a fixed old kit forever, and pinning it to `package.json` would request an unpublished version between the release-prep commit and the npm publish. The exemption covers `docks-kit` alone. Both installers still pin the Bun installer they download to the manifest's verified version, and `cli/test/unit/install.test.ts` asserts that pin in all four launcher and installer scripts.
110
120
  ## Testing
111
121
 
112
- Automated coverage includes `bun run test:unit`, `bun run golden:dryrun`, and `bun run golden:mutation`; prove-red modes must exit non-zero after detecting planted mismatches. Also verify user-facing changes via `./docks-kit sync --dry-run`, per-tool sanity (`/doctor`, `/plugin`, etc.), and `diff <(jq -S . <SoT>) <(jq -S . <deployed>)` recipes from the per-tool file.
122
+ Run `bun run check` before and after each change: oxlint, `oxfmt --check`, `tsc --noEmit`, unit tests, and both golden suites. Scripts: `lint`, `format:check`, `typecheck`, `test:unit`, `test:integration` (goldens), `test`, `format`; CI runs `check` on push and pull request with Bun 1.4.2. Prove-red modes must exit non-zero after detecting planted mismatches. The golden writer emits two-space JSON that `oxfmt` re-serializes, so run `bun run format` after `--update-goldens` or `format:check` fails. Also verify user-facing changes via `./docks-kit sync --dry-run`, per-tool sanity (`/doctor`, `/plugin`, etc.), `bun run test:runtime:posix`, and `diff <(jq -S . <SoT>) <(jq -S . <deployed>)` recipes from the per-tool file.
113
123
 
114
124
  Use direct acceptance and focused regressions while iterating, then run the full unit/golden gate once at the pre-commit or release boundary. Reuse still-matching evidence; a later relevant edit invalidates only the affected rung and final gate, not every prior check.
115
125
 
@@ -142,7 +152,7 @@ The Output Standard wins.
142
152
  Use `unslop` only for its pattern-detection lists.
143
153
 
144
154
  <constraint>
145
- When a kit-mechanic skill, its `references/`, or a wrapper agent (`.claude/agents/*.md` + its `.codex/agents/*.toml` twin) cites EngineNative internals, name the **module + exported/local function + semantic anchor** (e.g. `claudePlugins.ts syncPlugins, pass 5 uninstall guard`) — never a raw `file:NNN` line number, which goes stale on every refactor. Keep exactly one coarse `metadata.source_files[].lines` range per skill file as the sole intentional line-number touchpoint.
155
+ When a kit-mechanic skill, its `references/`, or a wrapper agent (`.claude/agents/*.md` + its `.codex/agents/*.toml` twin) cites EngineNative internals, name the **module + exported/local function + semantic anchor** (e.g. `claudePluginPasses.ts syncPlugins, pass 5 uninstall guard`) — never a raw `file:NNN` line number, which goes stale on every refactor. Keep exactly one coarse `metadata.source_files[].lines` range per skill file as the sole intentional line-number touchpoint.
146
156
  </constraint>
147
157
 
148
158
  **Universal-skill bootstrap.** `SoT/.agents/skills.txt` is intentionally empty, so the default sync exposes no universal skills. The generic opt-in contract remains: each [agentskills.io](https://agentskills.io/specification) slug is installed by `skillsSync.ts` with `npx skills add <slug> -g -y -a claude-code codex`. The slug comes first because `-a/--agent` is variadic; naming both supported agents preserves the canonical `~/.agents/skills/<name>/SKILL.md` plus Claude's `~/.claude/skills/<name>` symlink. Existing canonical directories are reused idempotently, `--prune` removes only entries tracked in `~/.agents/.kit-managed-skills`, and skills requiring a separate CLI retain explicit helpers in `skillsSync.ts`.
@@ -158,33 +168,31 @@ non-`local` effect.
158
168
 
159
169
  <constraint>
160
170
  The plan record is a GitHub issue. Its body starts with
161
- `<!-- plan-contract: v3 -->`, then a blank line and the exact eight `##`
162
- sections; it has no frontmatter. GitHub owns title, open-work phase, owner,
163
- timestamps, and completion, and no plan markdown is tracked in the repository.
164
- Exactly three skills own the workflow: `plan-workspace` maintains the workspace;
165
- main-context `plan-manager` runs six phases - decide, draft, research, plan
166
- review, implement, code review - with bounded repair and fresh re-review in both
167
- review phases, then archives; internal `plan-reviewer` returns one readable
168
- pre-implementation verdict block per round. Two read-only reviewer wrappers
169
- ship, `plan-reviewer` and `code-reviewer`, and nothing else in the lifecycle has
170
- a wrapper.
171
+ `<!-- plan-contract: v4 -->`, followed by one blank line and seven `##`
172
+ sections. It has no frontmatter. GitHub owns the title, open-work phase, owner,
173
+ timestamps, and completion. The repository tracks no plan markdown.
174
+ Exactly three skills own the workflow. `plan-workspace` maintains the workspace.
175
+ Main-context `plan-manager` runs six phases: decide, draft, research, plan
176
+ review, implement, and code review. It uses bounded repair and fresh re-review
177
+ in both review phases, then archives. Internal `plan-reviewer` returns one
178
+ readable pre-implementation verdict block per round. Two read-only reviewer
179
+ wrappers ship: `plan-reviewer` and `code-reviewer`.
171
180
  </constraint>
172
181
 
173
- After the marker and blank line, the record carries exactly `## Goal`,
174
- `## Research`, `## Steps`, `## Acceptance`, `## Do not touch`,
175
- `## Open questions`, `## Review`, and `## Verification Results`, in that order
176
- and once each. `## Goal` carries exactly one mode line. Open-work phase is one
177
- of `drafting`, `planned`, `ongoing`, or `blocked` in a `plan:<phase>` label; a
178
- blocked plan starts `## Open questions` with `Blocked: <one-line reason>`.
179
- Closed completion derives from GitHub `state` and `stateReason`. `## Review`
180
- contains exactly `_Review records are stored in issue comments._`. Each reviewer
181
- returns one markdown block, and the manager posts that whole block as one issue
182
- comment. The latest trusted well-formed record per review kind wins; its author
183
- must equal the plan's sole assignee. A legacy body verdict is consulted only
184
- when no trusted comment record exists for that kind. Both review phases use
185
- fresh inputs and run at most five rounds, stopping on pass, no progress, a
186
- finding surviving its fix, or `repair` or `fixes-required` in round five. A
187
- plan-review `blocked` verdict always routes its user-only decision through
182
+ After the marker and blank line, the record carries exactly seven sections.
183
+ They appear once each in this order: `## Goal`, `## Research`, `## Steps`,
184
+ `## Acceptance`, `## Do not touch`, `## Open questions`, and
185
+ `## Verification Results`. `## Goal` carries exactly one mode line.
186
+ Open-work phase is `drafting`, `planned`, `ongoing`, or `blocked`.
187
+ Every plan carries `plan`. A `plan:<phase>` label stores the open phase.
188
+ A blocked plan starts `## Open questions` with `Blocked: <one-line reason>`.
189
+ GitHub `state` and `stateReason` determine closed status. Review records live
190
+ only in issue comments. Each reviewer returns one markdown block.
191
+ The manager posts that block as one issue comment. The latest trusted eligible
192
+ comment per review kind wins. Its author must equal the plan's sole assignee.
193
+ Both review phases use fresh inputs and run at most five rounds. They stop on
194
+ pass, no progress, a finding that survives repair, or a round-five non-pass.
195
+ A plan-review `blocked` verdict routes its user-only decision through
188
196
  `## Open questions` and `ask`.
189
197
 
190
198
  The record carries no hash, permit, run identity, lock, or bundle, and the
@@ -0,0 +1,104 @@
1
+ # omp context and compaction
2
+
3
+ This topic records why `SoT/.omp/config.yml` does not pin a compaction
4
+ threshold, and what omp's reserve-based default computes on each model lane.
5
+
6
+ ## How omp resolves the trigger
7
+
8
+ `resolveThresholdTokens` in omp source
9
+ `packages/agent/src/compaction/compaction.ts` selects the trigger from three
10
+ branches, in priority order.
11
+
12
+ 1. A positive `compaction.thresholdTokens` wins. omp clamps it to the range 1
13
+ through `contextWindow - 1`.
14
+ 2. Otherwise a positive `compaction.thresholdPercent` applies. omp clamps the
15
+ percent to the range 1 through 99 and computes
16
+ `floor(contextWindow * percent / 100)`.
17
+ 3. Otherwise, when both keys hold `-1`, reserve-based behavior applies.
18
+
19
+ The reserve-based branch computes the trigger from the window itself:
20
+
21
+ ```text
22
+ threshold = contextWindow - max(floor(contextWindow * 0.15), 16384)
23
+ ```
24
+
25
+ omp names the 16384 constant `DEFAULT_RESERVE_TOKENS`. The 15% proportional
26
+ reserve dominates on any window above roughly 109,227 tokens, so every lane
27
+ the kit uses resolves through the proportional term and never through the flat
28
+ constant.
29
+
30
+ ## Per-lane result
31
+
32
+ | Lane | Window | Reserve | Threshold | Percent of window |
33
+ |---|---:|---:|---:|---:|
34
+ | `openai-codex/*` | 272,000 | 40,800 | 231,200 | 85.0% |
35
+ | `anthropic/*` | 1,000,000 | 150,000 | 850,000 | 85.0% |
36
+
37
+ Both lanes compact at the same fraction of their own window. The absolute
38
+ trigger differs only because the windows differ.
39
+
40
+ ## Why the kit no longer pins 231200
41
+
42
+ The kit previously shipped `compaction.thresholdTokens: 231200`. That value is
43
+ exactly the reserve-based default for a 272,000-token window. It was therefore
44
+ identical to the default on the Codex lane and changed nothing there. Its only
45
+ real effect was forcing the Anthropic lane to compact at 23.1% of its window
46
+ instead of 85%.
47
+
48
+ The evidence below records 83 compactions measured between 2026-09-12 and
49
+ 2026-09-21, after omp 18.1.18 introduced the Anthropic server-side compaction
50
+ lane.
51
+
52
+ | Lane | Compactions | Maximum tokens before | Over window | Warnings |
53
+ |---|---:|---:|---:|---:|
54
+ | `anthropic` | 75 | 248,583 (24.9% of window) | 0 | 0 |
55
+ | `openai-codex` | 8 | 263,591 (96.9% of window) | 0 | 0 |
56
+
57
+ The median context freed was 64.8%. The pin was safe rather than dangerous.
58
+ No session ran past its window, and omp raised no warning on either lane.
59
+ Under the reserve-based default none of the 75 Anthropic compactions would
60
+ have fired, because no session reached 850,000.
61
+
62
+ ## Why the key ships as -1 rather than being deleted
63
+
64
+ `ompYaml.ts` `mergeOmpConfig` is additive. Its `mergeMappings, deployed-key
65
+ retention loop` returns every deployed key that the SoT omits to the merged
66
+ output, and the only drop predicate is `dropFallbackWildcard`, which is scoped
67
+ to `retry.fallbackChains`.
68
+
69
+ Deleting the key from the SoT would therefore leave the stale `231200` in
70
+ every already-deployed `~/.omp/agent/config.yml` permanently. A present key
71
+ holding the schema default sentinel `-1` overwrites that stale value on the
72
+ next sync.
73
+
74
+ Verify the deployed value after a sync:
75
+
76
+ ```bash
77
+ omp config get compaction.thresholdTokens
78
+ ```
79
+
80
+ The command must print `-1`.
81
+
82
+ ## The clamp hazard a fixed pin carries
83
+
84
+ The fixed-token branch clamps the pinned value to `contextWindow - 1`. On any
85
+ model whose window is below the pinned value, the reserve collapses to one
86
+ token, so omp compacts only once the context is already full. The
87
+ reserve-based default always leaves at least 15% of the window free.
88
+
89
+ No current role model has a window below 272,000, so the hazard was latent
90
+ rather than live. Not pinning removes it entirely, including for any future
91
+ model with a smaller window.
92
+
93
+ ## Related settings left at their omp defaults
94
+
95
+ | Setting | Value | Reason |
96
+ |---|---|---|
97
+ | `extendedContext` | `false` | omp 17.4.0 added the setting defaulting on, and a later release flipped the default to off. It caps a model carrying a premium long-context price tier at that model's standard-pricing window, so `openai-codex/gpt-5.6-sol` reports 272,000 rather than about 1,050,000. Enabling it makes a Codex window overrun structurally impossible. The cost is a cliff rather than a marginal rate: any request above the 272,000 threshold bills the whole request at input $10 instead of $4, cache read $1.00 instead of $0.40, and cache write $12.50 instead of $5.00. |
98
+ | `compaction.idleEnabled` | `true` | omp default. Idle compaction fires below the reserve-based trigger on both lanes. A compaction whose reason is `idle` skips the progress guard. |
99
+ | `compaction.idleThresholdTokens` | `200000` | omp default. |
100
+ | `compaction.idleTimeoutSeconds` | `300` | omp default. |
101
+ | `compaction.methodOrder` | `["remote","snapcompact","handoff","shake","soft"]` | omp default, remote first. Since omp 18.1.18 the Anthropic server-side lane handles Anthropic models, so snapcompact is a fallback rather than the primary path. |
102
+
103
+ Every row records a decision to leave the omp default in place. None of these
104
+ keys carries a kit value that differs from the default omp already applies.
@@ -11,7 +11,7 @@ against.
11
11
  | `default` | `anthropic/claude-opus-5` | high | 48 | $3.61 | 16.96 s | 66 (Claude Code) |
12
12
  | `slow` | `anthropic/claude-opus-5` | xhigh | 50 | $4.88 | 28.65 s | 68 (Claude Code) |
13
13
  | `plan` | `anthropic/claude-opus-5` | xhigh | 50 | $4.88 | 28.65 s | 68 (Claude Code) |
14
- | `task` | `openai-codex/gpt-6-astra` | low | 46 | $0.82 | 2.60 s | n/a |
14
+ | `task` | `openai-codex/gpt-5.6-sol` | high | 42 | $0.81 | 11.26 s | 64 (Codex) |
15
15
  | `advisor` | `openai-codex/gpt-5.6-sol` | medium | 39 | $0.50 | 4.90 s | 62 (Codex) |
16
16
  | `designer` | `anthropic/claude-opus-5` | high | 48 | $3.61 | 16.96 s | 66 (Claude Code) |
17
17
  | `vision` | `anthropic/claude-opus-5` | medium | 45 | $2.19 | 3.79 s | 64 (Claude Code) |
@@ -19,17 +19,19 @@ against.
19
19
  | `tiny` | `openai-codex/gpt-5.6-luna` | low | 22 | n/a | 1.78 s | 25 (Codex) |
20
20
  | `fable` | `anthropic/claude-fable-5-1` | medium | 49 | $2.98 | 9.65 s | n/a |
21
21
  | `switch_fable` | `anthropic/claude-fable-5-1` | medium | 49 | $2.98 | 9.65 s | n/a |
22
+ | `astra` | `openai-codex/gpt-6-astra` | xhigh | 53 | $2.31 | 161.65 s | n/a |
23
+ | `web` | `web/firecrawl` | n/a | n/a | n/a | n/a | n/a |
22
24
 
23
25
  The table reports the measured Artificial Analysis figures for each assigned
24
26
  model and level. It states no motive that the config or omp's own
25
27
  documentation does not establish. AA lists no cost per task for Luna medium
26
28
  and low, and no Coding Agent Index for any Astra or Fable 5.1 level except
27
- max.
29
+ max. AA measures no web search provider, so the `web` row carries no figures.
28
30
 
29
31
  What omp's settings catalog establishes about these roles:
30
32
 
31
- - `cycleOrder` lists the roles the model switcher cycles, so `fable` is the
32
- fourth `Ctrl+P` stop.
33
+ - `cycleOrder` lists the roles the model switcher cycles. `fable` is the
34
+ fourth `Ctrl+P` stop, and `astra` is the fifth and last stop.
33
35
  - `tiny` overrides the model for lightweight background tasks: titles, memory,
34
36
  auto-thinking, and unexpected-stop detection.
35
37
  - `modelTags` carries role metadata and can introduce roles; `hidden: true`
@@ -44,6 +46,28 @@ are dormant until an OMP agent with that name exists.
44
46
  `retry.fallbackChains.task` keeps `anthropic/claude-opus-5:high` as a
45
47
  cross-vendor fallback.
46
48
 
49
+ `retry.fallbackChains.astra` holds `anthropic/claude-fable-5-1:medium`.
50
+ `retry.fallbackChains.fable` holds `openai-codex/gpt-6-astra:xhigh`.
51
+ Each deliberate cycle stop falls to the other vendor. Without these explicit
52
+ chains, `retry.fallbackChains.default` would send either stop to
53
+ `openai-codex/gpt-5.6-sol:high`.
54
+ Chain entries are concrete selectors, not role aliases, so this pair cannot
55
+ recurse. The hidden `switch_fable` chain stays empty.
56
+
57
+ `modelRoles.web` is `web/firecrawl`, and `retry.fallbackChains.web` lists the
58
+ explicit 27-entry provider order that follows it. The two keys replace the
59
+ retired `providers.webSearchOrder` key, which omp no longer carries in its
60
+ settings schema. omp still accepts that key in a deployed file, expands it in
61
+ memory into the same two keys, and then drops it, but it never writes the
62
+ expansion back to disk. The kit therefore declares both keys itself.
63
+
64
+ The chain is the verbatim expansion omp produces today, read back with
65
+ `omp config get retry.fallbackChains`. It must stay complete. An explicit
66
+ chain replaces omp's built-in web order wholesale, so a shortened list drops
67
+ providers instead of reordering them. The first five entries keep the previous
68
+ Firecrawl, Exa, Perplexity, Gemini, Codex preference; the remaining entries
69
+ are omp's own ordering of the providers behind it.
70
+
47
71
  ## Artificial Analysis snapshot
48
72
 
49
73
  Source: `https://artificialanalysis.ai`, read on 2026-09-08. Every score below
@@ -139,9 +163,14 @@ published numeric cache price. Context 1M.
139
163
  Price: $0.20 in, $1.20 out, $0.02 cache read per 1M; no published cache-write
140
164
  price. Context 1M. AA lists cost per task only for max and xhigh.
141
165
 
142
- ## Why `task` points at Astra low
166
+ ## Why `task` returns to Sol high
167
+
168
+ `task` runs `openai-codex/gpt-5.6-sol:high`. The owner uses Astra only for
169
+ main orchestration, so Astra now has a dedicated `astra` cycle stop at `xhigh`.
170
+ The comparison below records the retired Astra-low choice beside the current
171
+ Sol-high choice and the other measured alternatives.
143
172
 
144
- | Metric | Astra low | Sol high | Sol max | Opus 5 high |
173
+ | Metric | Astra low, retired task | Sol high, current task | Sol max | Opus 5 high |
145
174
  |---|---:|---:|---:|---:|
146
175
  | Intelligence Index | 46 | 42 | 47 | 48 |
147
176
  | Cost per Index task | $0.82 | $0.81 | $1.99 | $3.61 |
@@ -154,58 +183,53 @@ price. Context 1M. AA lists cost per task only for max and xhigh.
154
183
  | AA-Briefcase | 1253 | 1361 | 1475 | 1557 |
155
184
  | AA-Omniscience | 41 | 20 | 22 | 34 |
156
185
 
157
- Astra low beats the previous `task` model, Sol high, on intelligence, latency,
158
- and token use at the same cost per task. Astra costs 2.5× per token and spends
159
- about one third the tokens, so the price rise and the efficiency gain cancel:
160
- this is a latency and token-budget win, not a cost saving.
161
-
162
- Sol medium is the cheaper measured alternative, and it was not chosen: index
163
- 39, Coding Agent Index 62 (Codex), $0.50 per task, 4.90 s TTFT. Against it,
164
- Astra low costs 64% more per task and scores 7 index points higher, with no
165
- published Astra coding-agent score at that level.
166
-
167
- Two risks come with it. Astra low loses AA-Briefcase, the eval closest to this
168
- kit's agent workload, and AA publishes no Coding Agent Index score for any Astra
169
- level except max. Watch reviewer output, because `reviewer` and
170
- `security-reviewer` inherit `@task`. If review quality drops, pin those agents
171
- to `anthropic/claude-opus-5:high` rather than reverting the whole role.
172
-
173
- ## What resolves to Astra low in practice
174
-
175
- `modelRoles.task` sets the model. The `:low` suffix alone does not cap every
176
- spawn. `SoT/.omp/models.yml` restricts the `openai-codex/gpt-6-astra` effort
177
- ladder to `[low]` through
178
- `providers.openai-codex.modelOverrides.gpt-6-astra.thinking`. omp clamps any
179
- requested effort to the model ladder, and `auto` has only one choice. Every
180
- Astra spawn therefore runs `low`, regardless of its `effort` hint.
186
+ Astra low beat Sol high on intelligence, latency, and token use at effectively
187
+ equal cost per task. It cost $0.82 against $0.81 and used 4k output tokens
188
+ per task against 13k. Its 2.60 s TTFT beat Sol high's 11.26 s.
189
+ That was a latency and token-budget win, not a cost saving.
190
+ The owner reversed that trade on purpose to reserve Astra for interactive
191
+ orchestration.
192
+
193
+ Sol high scores 64 in the Codex Coding Agent Index and 1361 on AA-Briefcase.
194
+ Astra low scores 1253 on AA-Briefcase. AA publishes no Astra Coding Agent
195
+ Index except max, which scores 67 in Codex.
196
+ The new Astra xhigh cycle stop scores 53 on the Intelligence Index, costs
197
+ $2.31 per index task, and has 161.65 s TTFT.
198
+ It has no published Coding Agent Index.
199
+
200
+ ## How Astra is selected in practice
201
+
202
+ No subagent role or `task.agentModelOverrides` entry resolves Astra.
203
+ `modelRoles.astra` is reachable only through the model switcher as a main
204
+ orchestrator role. The explicit Fable retry chain can also select Astra.
205
+ No subagent selects Astra automatically.
206
+
207
+ `SoT/.omp/models.yml` declares the full `low, medium, high, xhigh, max` ladder
208
+ with `defaultLevel: xhigh` under
209
+ `providers.openai-codex.modelOverrides.gpt-6-astra.thinking`.
210
+ The override stays as the kit's worked example of a provider ladder
211
+ redefinition. The in-session thinking control can still select `low` for a
212
+ quick answer.
181
213
 
182
214
  - The bundled `scout` and `sonic` agents carry `model: "@smol"` and
183
215
  `thinking-level: medium` in their embedded frontmatter, so they run Luna,
184
- not Astra. To move them, change `modelRoles.smol` or add a
216
+ not Sol or Astra. To move them, change `modelRoles.smol` or add a
185
217
  `task.agentModelOverrides` entry for the agent name.
186
218
  - The bundled `task` agent carries `model: "@task"` and
187
- `thinking-level: auto`. `auto` classifies each prompt and picks a level for
188
- the resolved model, but Astra's ladder permits only `low`.
219
+ `thinking-level: auto`. It resolves Sol, and `auto` classifies each prompt
220
+ to choose a thinking level.
189
221
  - `task.enableEffort` is `true`, so a caller can pass `effort: lo`, `med`, or
190
222
  `hi`, which overrides `auto`.
191
- - `task.maxEffort` is `high`, so `scout` and `sonic` run Luna `medium` by
192
- default and Luna `high` with `effort: hi`.
223
+ - `task.maxEffort` is `max`, so `scout` and `sonic` run Luna `medium` by
224
+ default and Luna `max` with `effort: hi`.
225
+ - The bundled `reviewer` and `security-reviewer` inherit `@task`, now Sol high.
193
226
  - The `code-reviewer` and `plan-reviewer` override entries remain dormant.
194
- omp's task tool rejects both names as unknown agents, so neither can spawn.
195
-
196
- Fresh `omp -p` runs on 2026-09-09 verified the model ladder with
197
- `task.maxEffort: high`. Child session logs recorded these results:
227
+ Both point to `@task`, now Sol high. omp's task tool rejects both names as
228
+ unknown agents, so neither can spawn.
198
229
 
199
- | Agent | Effort hint | Resolved model and level |
200
- |---|---|---|
201
- | `task` | `hi` | `gpt-6-astra:low` |
202
- | `task` | None, complex prompt | `gpt-6-astra:low` |
203
- | `reviewer` | `hi` | `gpt-6-astra:low` |
204
- | `security-reviewer` | None | `gpt-6-astra:low` |
205
- | `scout` | `hi` | `gpt-5.6-luna:high` |
206
- | `sonic` | `hi` | `gpt-5.6-luna:high` |
207
-
208
- The complex `task` run completed with 682 output tokens.
230
+ The runtime per-agent measurement from fresh `omp -p` runs on 2026-09-09
231
+ predates this change. It applies to the retired Astra-low `task` configuration,
232
+ not the current role map.
209
233
 
210
234
  omp's 272k context window is the `/extended-context off` window for Astra.
211
235
  `/extended-context on` uses the 922k input window, 1.05M total, which matches
@@ -214,8 +238,9 @@ AA's 1M.
214
238
  ## Free session launcher (`docks-kit omp`)
215
239
 
216
240
  `docks-kit omp [--model <selector>|--pick] [args...]` starts one interactive
217
- omp session on a single free model. All 12 model roles resolve to that model.
218
- All 9 retry fallback chains are empty, so a retry cannot reach a paid model.
241
+ omp session with a configuration overlay that selects a free model. The
242
+ overlay sets all 13 model roles to that model and empties all 10 retry fallback
243
+ chains. Higher-precedence model selection can replace those values.
219
244
  The overlay also sets `defaultThinkingLevel` and `task.maxEffort`, except for
220
245
  a model that publishes no thinking ladder, where both keys are omitted and the
221
246
  deployed values apply.
@@ -235,6 +260,28 @@ Every argument after the launcher flags forwards verbatim to omp. `docks-kit
235
260
  omp -p "..."` runs one prompt. `docks-kit omp --continue` resumes the
236
261
  previous session. A bare `docks-kit omp` opens an interactive session.
237
262
 
263
+ The overlay ranks above the deployed global and project configuration, and
264
+ below runtime overrides and later overlay files. Each input below therefore
265
+ selects a model that the overlay does not control:
266
+
267
+ - A forwarded later config overlay. The launcher passes its own `--config`
268
+ first, and omp applies a later overlay file over an earlier one.
269
+ - The runtime model flags `--model`, `--smol`, `--slow`, and `--plan`. The
270
+ legacy `--provider` flag selects a provider. `--models` sets the model
271
+ patterns that `Ctrl+P` cycling can reach, so cycling can leave the free
272
+ model.
273
+ - The model environment variables `PI_SMOL_MODEL`, `PI_SLOW_MODEL`, and
274
+ `PI_PLAN_MODEL`. `PI_CONFIG_FILES` is not in this set, because those files
275
+ load before `--config` overlays.
276
+ - The in-session model selector. `Alt+M` sets the roles, and `Alt+P` picks a
277
+ model for the current session only.
278
+ - A restored persisted session model. `--continue`, `--resume`, and
279
+ `autoResume` restore the model of the session they open, so a session that
280
+ started under the paid configuration returns to its paid model.
281
+
282
+ This list names the known selection paths. It is not a closed set. The
283
+ launcher does not reject these inputs today.
284
+
238
285
  The default model is `opencode-zen/muse-spark-1.3-contributor-free` ("Muse
239
286
  Spark 1.3 Free", 1,048,576-token context) at `xhigh`. `xhigh` is the ceiling
240
287
  of that model ladder: `minimal, low, medium, high, xhigh`. The paid
@@ -258,9 +305,11 @@ committed. `--model <selector>` records one selector and derives both levels
258
305
  from the catalog row. `--pick` opens an interactive wizard.
259
306
 
260
307
  The picker lists only models the live `omp models --json` catalog reports at
261
- zero input and output cost (26 entries on 2026-09-11). This path cannot start
262
- a paid session. The catalog can advertise a free model that the account
263
- cannot call; omp reports that provider error unchanged.
308
+ zero input and output cost (26 entries on 2026-09-11). The resulting overlay
309
+ sets every role to the chosen free model and empties every retry chain.
310
+ Higher-precedence model selection can replace those values. The catalog can
311
+ advertise a free model that the account cannot call; omp reports that provider
312
+ error unchanged.
264
313
 
265
314
  After the model, the wizard asks about thinking levels. `ompOverlay.ts
266
315
  planEffortChoice` decides which questions apply, because omp accepts a
@@ -299,5 +348,6 @@ recorded default and the ladder counts in these docs in the same commit.
299
348
  - Record the index version with the numbers. AA changes index composition
300
349
  between versions, so a score from another version is not a comparison.
301
350
  - Update the capture date in the same commit as any number.
302
- - Verify changes to the Astra ladder in `SoT/.omp/models.yml` and to
303
- `task.maxEffort` together with a fresh `omp -p` spawn per bundled agent.
351
+ - Verify Astra ladder changes in `SoT/.omp/models.yml` through the model
352
+ switcher and in-session thinking control.
353
+ - Verify `task.maxEffort` changes with a fresh `omp -p` spawn per bundled agent.