docks-kit 0.17.0 → 0.17.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (84) hide show
  1. package/AGENTS.md +37 -33
  2. package/README.md +1 -0
  3. package/cli/docs/omp-models.md +114 -58
  4. package/cli/src/argv.ts +38 -469
  5. package/cli/src/argvSurface.ts +371 -0
  6. package/cli/src/argvValidate.ts +122 -0
  7. package/cli/src/commands/docs.ts +57 -42
  8. package/cli/src/commands/harnesses.ts +38 -45
  9. package/cli/src/commands/model.ts +48 -53
  10. package/cli/src/commands/models.ts +24 -26
  11. package/cli/src/commands/omp.ts +178 -88
  12. package/cli/src/commands/plugins.ts +20 -20
  13. package/cli/src/commands/skills.ts +22 -20
  14. package/cli/src/commands/status.ts +87 -80
  15. package/cli/src/commands/sync.ts +113 -87
  16. package/cli/src/commands/toolchain.ts +21 -21
  17. package/cli/src/commands/update.ts +139 -117
  18. package/cli/src/efforts.ts +35 -34
  19. package/cli/src/engine-native/DESIGN.md +30 -30
  20. package/cli/src/engine-native/bun.ts +77 -55
  21. package/cli/src/engine-native/claudeLsp.ts +154 -0
  22. package/cli/src/engine-native/claudeOptionalPlugins.ts +113 -0
  23. package/cli/src/engine-native/claudePluginPasses.ts +348 -0
  24. package/cli/src/engine-native/claudeRemovals.ts +195 -0
  25. package/cli/src/engine-native/claudeRetired.ts +4 -4
  26. package/cli/src/engine-native/claudeRuntime.ts +94 -68
  27. package/cli/src/engine-native/claudeSettings.ts +174 -0
  28. package/cli/src/engine-native/claudeSettingsModifiers.ts +50 -54
  29. package/cli/src/engine-native/claudeSync.ts +176 -456
  30. package/cli/src/engine-native/codexConfig.ts +158 -0
  31. package/cli/src/engine-native/codexHooks.ts +91 -0
  32. package/cli/src/engine-native/codexPlugins.ts +290 -0
  33. package/cli/src/engine-native/codexStatus.ts +24 -0
  34. package/cli/src/engine-native/codexSync.ts +110 -568
  35. package/cli/src/engine-native/codexToml.ts +140 -138
  36. package/cli/src/engine-native/deps.ts +144 -107
  37. package/cli/src/engine-native/engineCtx.ts +126 -0
  38. package/cli/src/engine-native/exec.ts +98 -87
  39. package/cli/src/engine-native/failures.ts +3 -3
  40. package/cli/src/engine-native/harnesses.ts +62 -72
  41. package/cli/src/engine-native/index.ts +34 -315
  42. package/cli/src/engine-native/jq.ts +25 -24
  43. package/cli/src/engine-native/logger.ts +98 -100
  44. package/cli/src/engine-native/models.ts +50 -45
  45. package/cli/src/engine-native/modes.ts +83 -80
  46. package/cli/src/engine-native/ompFileDeploy.ts +111 -0
  47. package/cli/src/engine-native/ompMarketplace.ts +136 -0
  48. package/cli/src/engine-native/ompOverlay.ts +93 -53
  49. package/cli/src/engine-native/ompPaths.ts +49 -38
  50. package/cli/src/engine-native/ompPlugins.ts +170 -0
  51. package/cli/src/engine-native/ompSync.ts +64 -408
  52. package/cli/src/engine-native/ompYaml.ts +43 -41
  53. package/cli/src/engine-native/os/darwin.ts +5 -5
  54. package/cli/src/engine-native/os/index.ts +15 -15
  55. package/cli/src/engine-native/os/linux.ts +5 -5
  56. package/cli/src/engine-native/os/posix.ts +28 -24
  57. package/cli/src/engine-native/os/targets.ts +17 -17
  58. package/cli/src/engine-native/os/types.ts +31 -31
  59. package/cli/src/engine-native/os/windows.ts +61 -61
  60. package/cli/src/engine-native/parseArgs.ts +122 -372
  61. package/cli/src/engine-native/parseHelp.ts +66 -0
  62. package/cli/src/engine-native/parseModifiers.ts +222 -0
  63. package/cli/src/engine-native/services.ts +34 -34
  64. package/cli/src/engine-native/settings.ts +15 -15
  65. package/cli/src/engine-native/sharedTypes.d.ts +46 -0
  66. package/cli/src/engine-native/skillsInstall.ts +105 -0
  67. package/cli/src/engine-native/skillsLinks.ts +203 -0
  68. package/cli/src/engine-native/skillsManifest.ts +16 -0
  69. package/cli/src/engine-native/skillsPrune.ts +108 -0
  70. package/cli/src/engine-native/skillsSync.ts +28 -362
  71. package/cli/src/engine-native/syncConcurrency.ts +55 -0
  72. package/cli/src/engine-native/syncDispatch.ts +141 -0
  73. package/cli/src/engine-native/toolchain.ts +65 -62
  74. package/cli/src/engine.ts +65 -57
  75. package/cli/src/generated/sotPayload.ts +7 -7
  76. package/cli/src/kitHome.ts +44 -40
  77. package/cli/src/main.ts +63 -56
  78. package/cli/src/manifests.ts +53 -54
  79. package/cli/src/md.d.ts +2 -2
  80. package/cli/src/payload.ts +9 -9
  81. package/cli/src/services.ts +21 -13
  82. package/cli/tsconfig.json +4 -1
  83. package/package.json +9 -2
  84. package/cli/src/engine-native/claudePlugins.ts +0 -510
package/AGENTS.md CHANGED
@@ -33,6 +33,10 @@ launcher can fall back to Bun source.
33
33
  | `cli/src/engine-native/ompSync.ts` | omp file deployment, marketplace registration, and plugin synchronization |
34
34
  | `cli/src/commands/omp.ts` | `docks-kit omp` session launcher: renders the free-model run overlay and forwards args to omp |
35
35
  | `cli/src/engine-native/ompOverlay.ts` | Free-model overlay render plus catalog parse and thinking-ceiling helpers |
36
+ | `cli/src/engine-native/<axis>.ts` | One change axis per file: `codexConfig`/`codexHooks`/`codexPlugins`/`codexStatus`, `claudeSettings`/`claudeRemovals`/`claudePluginPasses`/`claudeOptionalPlugins`/`claudeLsp`, `parseModifiers`/`parseHelp`, `ompFileDeploy`/`ompMarketplace`/`ompPlugins`, `skillsManifest`/`skillsLinks`/`skillsInstall`/`skillsPrune`, `engineCtx`/`syncConcurrency`/`syncDispatch`; orchestrators keep prior exports |
37
+ | `cli/src/engine-native/sharedTypes.d.ts` | Single definition per shared shape (manifest records, settings edits, omp session model, scalar modifier flags) |
38
+ | `.oxlintrc.json` / `.oxfmtrc.json` | Mechanical lint (correctness, suspicious, eqeqeq, no-var, prefer-const, no-unused) and code format scope |
39
+ | `cli/test/fixtures/` | Golden fixtures: `home-fresh`, `home-drift`, `home-invalid-json`, `codex-toml`, `statusline` |
36
40
  | `cli/` | Effect 4 RC CLI + bundled docs topics |
37
41
  | `SoT/models.json` | Kit-verified Claude and Codex model catalog |
38
42
  | `SoT/toolchain.json` | Toolchain floors manifest (verified pins consumed by EngineNative) |
@@ -66,16 +70,16 @@ omp SoT notes:
66
70
  - `SoT/.omp/AGENTS.md`, `config.yml`, `models.yml`, and `mcp.json` deploy to `~/.omp/agent/`.
67
71
  - `SoT/.omp/intercom.json` deploys to `$PI_CODING_AGENT_DIR/intercom/config.json`. The default root is `~/.pi/agent`.
68
72
  - `ompSync.ts syncMergedYaml` deep-merges `config.yml` through `ompYaml.ts mergeOmpConfig` and `models.yml` through `mergeOmpModels`. Both wrap one generic mapping merge; only the config wrapper prunes stale `retry.fallbackChains` wildcards.
69
- - `cycleOrder` ends with the `fable` role (`modelRoles.fable` = `anthropic/claude-fable-5-1:medium`, `modelTags.fable` visible, `retry.fallbackChains.fable` empty), so the model switcher reaches Fable 5.1 as its fourth stop and never falls back off it. The hidden `switch_fable` role points at the same model and level.
70
- - `modelRoles.task` is `openai-codex/gpt-6-astra:low`, chosen for 2.60 s TTFT and 4k output tokens per task at cost parity with the previous `gpt-5.6-sol:high`. `task.agentModelOverrides` points four reviewer names at `@task`; only the bundled `reviewer` and `security-reviewer` are discoverable OMP agents, so those two inherit `task`, and the `code-reviewer` and `plan-reviewer` entries are dormant. Artificial Analysis publishes no per-level Astra Coding Agent Index score, so reviewer output is the signal to watch.
71
- - `SoT/.omp/models.yml` deploys to `~/.omp/agent/models.yml` through `mergeOmpModels`, which keeps every deployed-only key. A user file may carry provider credentials, so whole-file replacement is wrong. It restricts the `gpt-6-astra` effort ladder to `low`, so every Astra subagent runs `low` while `task.maxEffort: high` lets `scout` and `sonic` reach Luna `high` with `effort: hi`. `cli/docs/omp-models.md` records the verification.
73
+ - `cycleOrder` ends with `astra` as its fifth stop. `modelRoles.astra` is `openai-codex/gpt-6-astra:xhigh`, and `modelTags.astra` is visible. Astra and Fable fall back to each other through concrete selectors. `modelRoles.fable` remains `anthropic/claude-fable-5-1:medium`, with visible `modelTags.fable`. The hidden `switch_fable` role uses the same Fable selector and keeps an empty fallback chain.
74
+ - `modelRoles.task` is `openai-codex/gpt-5.6-sol:high`. The four reviewer entries in `task.agentModelOverrides` inherit it through `@task`. Only bundled `reviewer` and `security-reviewer` are discoverable OMP agents; `code-reviewer` and `plan-reviewer` stay dormant.
75
+ - `SoT/.omp/models.yml` declares Astra's full `low, medium, high, xhigh, max` ladder with `defaultLevel: xhigh` as the worked provider ladder-override example. `ompYaml.ts mergeOmpModels` preserves deployed-only keys in `~/.omp/agent/models.yml`, because a user file may carry provider credentials. Whole-file replacement is wrong.
72
76
  - `SoT/.omp/AGENTS.md` carries the rule `Please remove all mannered prose.` Anthropic's Fable 5.1 prompting guide documents mannered prose as a Fable 5.1 behavior and gives that sentence as its short-version fix: https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1
73
77
  - `cli/docs/omp-models.md` (topic `omp-models`) records the role map rationale and the Artificial Analysis snapshot behind it. Model choices change with published benchmarks, so update that topic in the same commit as a role change.
74
78
  - Sync registers the `docks` marketplace. It installs or upgrades `docks@docks` and `plan-lifecycle@docks` at user scope.
75
79
  - Sync installs `pi-intercom` at the verified version from `SoT/toolchain.json`.
76
80
  - The omp CLI is upstream-owned and self-updating through `omp update`. Sync never installs or upgrades the CLI.
77
- - `docks-kit omp [--model <selector>|--pick] [args...]` starts one interactive omp session on a single free model. It renders a run overlay to `~/.cache/docks-kit/omp-free-<model>-<digest>.yml` (mode 0600) and passes it through omp's repeatable `--config` flag. `ompOverlay.ts overlayFileName` gives each model its own file, because omp can re-read the overlay during a live session and a second launcher on another model must not rewrite it. Deployed `~/.omp/agent/` files stay untouched, so the next plain `omp` run uses the paid configuration again. Remaining arguments forward verbatim to omp.
78
- - The session model persists per machine in `~/.docks-kit/state.json` under `ompSession`, next to `harnesses`. The default is `opencode-zen/muse-spark-1.3-contributor-free` at `xhigh`. `ompOverlay.ts ladderCeiling, advisorLevelFor` derive both the session level and the advisor level from the ladder of the chosen model, never from a fixed list: free ladders are not uniform, three free models publish no `medium`, and two publish no ladder, which the overlay renders as bare selectors. The picker lists only zero-cost catalog models.
81
+ - `docks-kit omp [--model <selector>|--pick] [args...]` starts one interactive omp session with every model role set to one free model and every retry chain empty. Higher-precedence model selection can replace those overlay values. It renders the overlay to `~/.cache/docks-kit/omp-free-<model>-<digest>.yml` (mode 0600) and passes it through omp's repeatable `--config` flag. `ompOverlay.ts overlayFileName` gives each model its own file, because omp can re-read the overlay during a live session and a second launcher on another model must not rewrite it. Deployed `~/.omp/agent/` files stay untouched, so the next plain `omp` run uses the paid configuration again. Remaining arguments forward verbatim to omp.
82
+ - The session model persists per machine in `~/.docks-kit/state.json` under `ompSession`, next to `harnesses`. The default is `opencode-zen/muse-spark-1.3-contributor-free` at `xhigh`. `ompOverlay.ts ladderCeiling, advisorLevelFor` derive both levels from the ladder of the chosen model for the `--model` path, and `ompOverlay.ts planEffortChoice` drives the `--pick` wizard, which asks whether every role shares one level and otherwise takes one level for the main roles and one for the advisor from that model's own ladder. Levels never come from a fixed list: free ladders are not uniform, three free models publish no `medium`, and two publish no ladder, which the overlay renders as bare selectors. The picker lists only zero-cost catalog models.
79
83
 
80
84
  For per-tool SoT layouts (`SoT/.claude/`, `SoT/.codex/`, `SoT/.omp/`), see the matching SoT directory.
81
85
 
@@ -90,8 +94,10 @@ For per-tool SoT layouts (`SoT/.claude/`, `SoT/.codex/`, `SoT/.omp/`), see the m
90
94
  - **Per-machine harness selection.** `~/.docks-kit/state.json` drives a flag-less sync. A missing file selects `claude`, `codex`, and `agents`; it never selects `omp` implicitly. `sync` never prompts and never writes the selection file. `docks-kit harnesses` is the only command that writes the `harnesses` selection. The same state file also carries the independent `ompSession` key, which only `docks-kit omp` writes; every writer merges over the stored record, so neither key can drop the other.
91
95
  - **Additive by default.** Keys present in deployed config but absent from SoT are preserved on default sync. This protects user-only additions, but means drift accumulates — neither flag-less reset can clean it up. The one exception is the Claude `removed` manifest (`claudeSync.ts REMOVED_MANIFEST, baseline removal inventory`), a curated list of unambiguous kit-owned artifacts that `claudeSync.ts syncRemovals, baseline artifact prune` force-prunes on every sync, including the home-relative `~/.local/bin/session-relay` artifact installed outside `~/.claude`; see `CLAUDE.md` § Pruning stale artifacts.
92
96
  - **`--reconcile` / `--prune` are the kit-owned reconcile flags.** Orthogonal — `--reconcile` reconciles the settings layer (SoT-declared keys/tables/arrays win; user-only keys and nested objects are preserved; permissions arrays are replaced wholesale by SoT). `--prune` uninstalls kit-managed installations not in the SoT (plugins, marketplaces, and `~/.agents/skills/*` entries tracked in `~/.agents/.kit-managed-skills`). Combine for a full reset to SoT's kit-managed scope. User-only additions outside the kit's scope (custom env vars, mcpServers, manually-installed skills, third-party plugins not declared in SoT) are always preserved. Each tool's per-tool file documents the specific paths and diff recipes.
93
- - **SOLID-aligned modules.** `cli/src/engine-native/parseArgs.ts` owns flag parsing and validation. `toolchain.ts` owns verified-version floor reporting over `SoT/toolchain.json`; `bun.ts` owns the shared, memoized Bun bootstrap; `claudeRuntime.ts` owns Claude settings materialization. `claudeSync.ts`, `codexSync.ts`, and `skillsSync.ts` own their tool-specific sync logic. `ompSync.ts` owns omp file and plugin sync. `ompYaml.ts` owns the omp YAML merge. `harnesses.ts` owns per-machine harness selection. `index.ts` is the thin orchestrator. The public CLI seam is `cli/src/engine.ts`.
97
+ - **SOLID-aligned modules.** Axis modules own one concern each (see layout); orchestrators (`claudeSync.ts`, `codexSync.ts`, `skillsSync.ts`, `ompSync.ts`, `index.ts`, `argv.ts`, `parseArgs.ts`) keep prior exports. New shared shapes go in `sharedTypes.d.ts`, never a second local copy. The public CLI seam is `cli/src/engine.ts`.
94
98
  - **Small, reviewable changes.** Bundled multi-concern PRs are harder to review and revert. Split an engine/CLI change and a per-tool config change unless the change requires atomicity.
99
+ - **Version bumps.** Bump `package.json` on user-visible behavior change; bump a `SoT/toolchain.json` `verified` pin only after testing that release. Refactors, splits, and doc edits bump nothing.
100
+ - **Regenerate the payload.** After `SoT/`, `notification.mp3`, `package.json`, or `cli/tsconfig.json` changes, run `bun cli/scripts/generate-sot-payload.ts`; `bun run check:generated` verifies it. After file moves, clamp cited `metadata.source_files[].lines` ranges so `skillMetadata.test.ts` passes.
95
101
  - **Dry-run before destructive flags.** Always preview with `./docks-kit sync --dry-run` (or the relevant `diff <(jq -S …)` recipe in the per-tool file) before invoking `--reconcile` or `--prune`. User-added permissions / env vars / plugins absent from SoT will be discarded.
96
102
  - **SoT prompt files are rules, not explanation.** `SoT/.claude/CLAUDE.md`, `SoT/.codex/AGENTS.md`, and `SoT/.omp/AGENTS.md` are loaded into agent sessions' prompt context — every line costs prompt tokens on every turn for every user. Restrict their content to rules, heuristics, and `<constraint>` blocks the agent must *act on* during a turn. Do NOT add inline source citations (`Source: …`, attributed quotes), "why this rule exists" preface text, version-watermarking trivia (e.g. "Distilled from X v2.0, captured 2025-11-07"), per-bug workarounds, or installation instructions. Provenance, motivation, and historical context belong in `CLAUDE.md` / `AGENTS.md` at the repo root (humans read once) or in commit messages — never in the SoT. For every line, apply the official test: …
97
103
  - **Cache-invariance for kit-authored prompt surfaces.** Never put timestamps, counters, or mutable state into SoT prompt files, hook outputs that land in the cached prefix, or tool definitions — cache breaks force cold-start writes. Dynamic context belongs in runtime-injected messages (e.g. SessionStart hook output), which is exactly how the kit's date/config injection works.
@@ -109,7 +115,7 @@ For per-tool SoT layouts (`SoT/.claude/`, `SoT/.codex/`, `SoT/.omp/`), see the m
109
115
  - **One exemption: the kit's own package.** `install.sh` and `install.ps1` end with `bun add -g docks-kit@latest`, because a global installer that pinned itself would install a fixed old kit forever, and pinning it to `package.json` would request an unpublished version between the release-prep commit and the npm publish. The exemption covers `docks-kit` alone. Both installers still pin the Bun installer they download to the manifest's verified version, and `cli/test/unit/install.test.ts` asserts that pin in all four launcher and installer scripts.
110
116
  ## Testing
111
117
 
112
- Automated coverage includes `bun run test:unit`, `bun run golden:dryrun`, and `bun run golden:mutation`; prove-red modes must exit non-zero after detecting planted mismatches. Also verify user-facing changes via `./docks-kit sync --dry-run`, per-tool sanity (`/doctor`, `/plugin`, etc.), and `diff <(jq -S . <SoT>) <(jq -S . <deployed>)` recipes from the per-tool file.
118
+ Run `bun run check` before and after each change: oxlint, `tsc --noEmit`, unit tests, and both golden suites. Scripts: `lint`, `typecheck`, `test:unit`, `test:integration` (goldens), `test`, `format`; CI runs `check` on push and pull request with Bun 1.4.2. Prove-red modes must exit non-zero after detecting planted mismatches. Also verify user-facing changes via `./docks-kit sync --dry-run`, per-tool sanity (`/doctor`, `/plugin`, etc.), `bun run test:runtime:posix`, and `diff <(jq -S . <SoT>) <(jq -S . <deployed>)` recipes from the per-tool file.
113
119
 
114
120
  Use direct acceptance and focused regressions while iterating, then run the full unit/golden gate once at the pre-commit or release boundary. Reuse still-matching evidence; a later relevant edit invalidates only the affected rung and final gate, not every prior check.
115
121
 
@@ -142,7 +148,7 @@ The Output Standard wins.
142
148
  Use `unslop` only for its pattern-detection lists.
143
149
 
144
150
  <constraint>
145
- When a kit-mechanic skill, its `references/`, or a wrapper agent (`.claude/agents/*.md` + its `.codex/agents/*.toml` twin) cites EngineNative internals, name the **module + exported/local function + semantic anchor** (e.g. `claudePlugins.ts syncPlugins, pass 5 uninstall guard`) — never a raw `file:NNN` line number, which goes stale on every refactor. Keep exactly one coarse `metadata.source_files[].lines` range per skill file as the sole intentional line-number touchpoint.
151
+ When a kit-mechanic skill, its `references/`, or a wrapper agent (`.claude/agents/*.md` + its `.codex/agents/*.toml` twin) cites EngineNative internals, name the **module + exported/local function + semantic anchor** (e.g. `claudePluginPasses.ts syncPlugins, pass 5 uninstall guard`) — never a raw `file:NNN` line number, which goes stale on every refactor. Keep exactly one coarse `metadata.source_files[].lines` range per skill file as the sole intentional line-number touchpoint.
146
152
  </constraint>
147
153
 
148
154
  **Universal-skill bootstrap.** `SoT/.agents/skills.txt` is intentionally empty, so the default sync exposes no universal skills. The generic opt-in contract remains: each [agentskills.io](https://agentskills.io/specification) slug is installed by `skillsSync.ts` with `npx skills add <slug> -g -y -a claude-code codex`. The slug comes first because `-a/--agent` is variadic; naming both supported agents preserves the canonical `~/.agents/skills/<name>/SKILL.md` plus Claude's `~/.claude/skills/<name>` symlink. Existing canonical directories are reused idempotently, `--prune` removes only entries tracked in `~/.agents/.kit-managed-skills`, and skills requiring a separate CLI retain explicit helpers in `skillsSync.ts`.
@@ -158,33 +164,31 @@ non-`local` effect.
158
164
 
159
165
  <constraint>
160
166
  The plan record is a GitHub issue. Its body starts with
161
- `<!-- plan-contract: v3 -->`, then a blank line and the exact eight `##`
162
- sections; it has no frontmatter. GitHub owns title, open-work phase, owner,
163
- timestamps, and completion, and no plan markdown is tracked in the repository.
164
- Exactly three skills own the workflow: `plan-workspace` maintains the workspace;
165
- main-context `plan-manager` runs six phases - decide, draft, research, plan
166
- review, implement, code review - with bounded repair and fresh re-review in both
167
- review phases, then archives; internal `plan-reviewer` returns one readable
168
- pre-implementation verdict block per round. Two read-only reviewer wrappers
169
- ship, `plan-reviewer` and `code-reviewer`, and nothing else in the lifecycle has
170
- a wrapper.
167
+ `<!-- plan-contract: v4 -->`, followed by one blank line and seven `##`
168
+ sections. It has no frontmatter. GitHub owns the title, open-work phase, owner,
169
+ timestamps, and completion. The repository tracks no plan markdown.
170
+ Exactly three skills own the workflow. `plan-workspace` maintains the workspace.
171
+ Main-context `plan-manager` runs six phases: decide, draft, research, plan
172
+ review, implement, and code review. It uses bounded repair and fresh re-review
173
+ in both review phases, then archives. Internal `plan-reviewer` returns one
174
+ readable pre-implementation verdict block per round. Two read-only reviewer
175
+ wrappers ship: `plan-reviewer` and `code-reviewer`.
171
176
  </constraint>
172
177
 
173
- After the marker and blank line, the record carries exactly `## Goal`,
174
- `## Research`, `## Steps`, `## Acceptance`, `## Do not touch`,
175
- `## Open questions`, `## Review`, and `## Verification Results`, in that order
176
- and once each. `## Goal` carries exactly one mode line. Open-work phase is one
177
- of `drafting`, `planned`, `ongoing`, or `blocked` in a `plan:<phase>` label; a
178
- blocked plan starts `## Open questions` with `Blocked: <one-line reason>`.
179
- Closed completion derives from GitHub `state` and `stateReason`. `## Review`
180
- contains exactly `_Review records are stored in issue comments._`. Each reviewer
181
- returns one markdown block, and the manager posts that whole block as one issue
182
- comment. The latest trusted well-formed record per review kind wins; its author
183
- must equal the plan's sole assignee. A legacy body verdict is consulted only
184
- when no trusted comment record exists for that kind. Both review phases use
185
- fresh inputs and run at most five rounds, stopping on pass, no progress, a
186
- finding surviving its fix, or `repair` or `fixes-required` in round five. A
187
- plan-review `blocked` verdict always routes its user-only decision through
178
+ After the marker and blank line, the record carries exactly seven sections.
179
+ They appear once each in this order: `## Goal`, `## Research`, `## Steps`,
180
+ `## Acceptance`, `## Do not touch`, `## Open questions`, and
181
+ `## Verification Results`. `## Goal` carries exactly one mode line.
182
+ Open-work phase is `drafting`, `planned`, `ongoing`, or `blocked`.
183
+ Every plan carries `plan`. A `plan:<phase>` label stores the open phase.
184
+ A blocked plan starts `## Open questions` with `Blocked: <one-line reason>`.
185
+ GitHub `state` and `stateReason` determine closed status. Review records live
186
+ only in issue comments. Each reviewer returns one markdown block.
187
+ The manager posts that block as one issue comment. The latest trusted eligible
188
+ comment per review kind wins. Its author must equal the plan's sole assignee.
189
+ Both review phases use fresh inputs and run at most five rounds. They stop on
190
+ pass, no progress, a finding that survives repair, or a round-five non-pass.
191
+ A plan-review `blocked` verdict routes its user-only decision through
188
192
  `## Open questions` and `ask`.
189
193
 
190
194
  The record carries no hash, permit, run identity, lock, or bundle, and the
package/README.md CHANGED
@@ -48,6 +48,7 @@ docks-kit harnesses view or change this machine's se
48
48
  docks-kit update [--no-sync] self-update the kit (autodetects checkout vs global install), then sync
49
49
  docks-kit model <claude|codex> [value] get/set the DEPLOYED model (TTY picker)
50
50
  docks-kit models [claude|codex] model catalogs (`--json`)
51
+ docks-kit omp [--model <m>|--pick] [args...] one omp session on a free model, nothing deployed changes
51
52
  docks-kit toolchain [check|ensure <tool>] verified-version floors for external tools
52
53
  docks-kit status [--json] deployed-vs-SoT drift + toolchain + counts
53
54
  docks-kit plugins list [--json] enabledPlugins tri-state vs installed
@@ -11,7 +11,7 @@ against.
11
11
  | `default` | `anthropic/claude-opus-5` | high | 48 | $3.61 | 16.96 s | 66 (Claude Code) |
12
12
  | `slow` | `anthropic/claude-opus-5` | xhigh | 50 | $4.88 | 28.65 s | 68 (Claude Code) |
13
13
  | `plan` | `anthropic/claude-opus-5` | xhigh | 50 | $4.88 | 28.65 s | 68 (Claude Code) |
14
- | `task` | `openai-codex/gpt-6-astra` | low | 46 | $0.82 | 2.60 s | n/a |
14
+ | `task` | `openai-codex/gpt-5.6-sol` | high | 42 | $0.81 | 11.26 s | 64 (Codex) |
15
15
  | `advisor` | `openai-codex/gpt-5.6-sol` | medium | 39 | $0.50 | 4.90 s | 62 (Codex) |
16
16
  | `designer` | `anthropic/claude-opus-5` | high | 48 | $3.61 | 16.96 s | 66 (Claude Code) |
17
17
  | `vision` | `anthropic/claude-opus-5` | medium | 45 | $2.19 | 3.79 s | 64 (Claude Code) |
@@ -19,6 +19,7 @@ against.
19
19
  | `tiny` | `openai-codex/gpt-5.6-luna` | low | 22 | n/a | 1.78 s | 25 (Codex) |
20
20
  | `fable` | `anthropic/claude-fable-5-1` | medium | 49 | $2.98 | 9.65 s | n/a |
21
21
  | `switch_fable` | `anthropic/claude-fable-5-1` | medium | 49 | $2.98 | 9.65 s | n/a |
22
+ | `astra` | `openai-codex/gpt-6-astra` | xhigh | 53 | $2.31 | 161.65 s | n/a |
22
23
 
23
24
  The table reports the measured Artificial Analysis figures for each assigned
24
25
  model and level. It states no motive that the config or omp's own
@@ -28,8 +29,8 @@ max.
28
29
 
29
30
  What omp's settings catalog establishes about these roles:
30
31
 
31
- - `cycleOrder` lists the roles the model switcher cycles, so `fable` is the
32
- fourth `Ctrl+P` stop.
32
+ - `cycleOrder` lists the roles the model switcher cycles. `fable` is the
33
+ fourth `Ctrl+P` stop, and `astra` is the fifth and last stop.
33
34
  - `tiny` overrides the model for lightweight background tasks: titles, memory,
34
35
  auto-thinking, and unexpected-stop detection.
35
36
  - `modelTags` carries role metadata and can introduce roles; `hidden: true`
@@ -44,6 +45,14 @@ are dormant until an OMP agent with that name exists.
44
45
  `retry.fallbackChains.task` keeps `anthropic/claude-opus-5:high` as a
45
46
  cross-vendor fallback.
46
47
 
48
+ `retry.fallbackChains.astra` holds `anthropic/claude-fable-5-1:medium`.
49
+ `retry.fallbackChains.fable` holds `openai-codex/gpt-6-astra:xhigh`.
50
+ Each deliberate cycle stop falls to the other vendor. Without these explicit
51
+ chains, `retry.fallbackChains.default` would send either stop to
52
+ `openai-codex/gpt-5.6-sol:high`.
53
+ Chain entries are concrete selectors, not role aliases, so this pair cannot
54
+ recurse. The hidden `switch_fable` chain stays empty.
55
+
47
56
  ## Artificial Analysis snapshot
48
57
 
49
58
  Source: `https://artificialanalysis.ai`, read on 2026-09-08. Every score below
@@ -139,9 +148,14 @@ published numeric cache price. Context 1M.
139
148
  Price: $0.20 in, $1.20 out, $0.02 cache read per 1M; no published cache-write
140
149
  price. Context 1M. AA lists cost per task only for max and xhigh.
141
150
 
142
- ## Why `task` points at Astra low
151
+ ## Why `task` returns to Sol high
152
+
153
+ `task` runs `openai-codex/gpt-5.6-sol:high`. The owner uses Astra only for
154
+ main orchestration, so Astra now has a dedicated `astra` cycle stop at `xhigh`.
155
+ The comparison below records the retired Astra-low choice beside the current
156
+ Sol-high choice and the other measured alternatives.
143
157
 
144
- | Metric | Astra low | Sol high | Sol max | Opus 5 high |
158
+ | Metric | Astra low, retired task | Sol high, current task | Sol max | Opus 5 high |
145
159
  |---|---:|---:|---:|---:|
146
160
  | Intelligence Index | 46 | 42 | 47 | 48 |
147
161
  | Cost per Index task | $0.82 | $0.81 | $1.99 | $3.61 |
@@ -154,58 +168,53 @@ price. Context 1M. AA lists cost per task only for max and xhigh.
154
168
  | AA-Briefcase | 1253 | 1361 | 1475 | 1557 |
155
169
  | AA-Omniscience | 41 | 20 | 22 | 34 |
156
170
 
157
- Astra low beats the previous `task` model, Sol high, on intelligence, latency,
158
- and token use at the same cost per task. Astra costs 2.5× per token and spends
159
- about one third the tokens, so the price rise and the efficiency gain cancel:
160
- this is a latency and token-budget win, not a cost saving.
161
-
162
- Sol medium is the cheaper measured alternative, and it was not chosen: index
163
- 39, Coding Agent Index 62 (Codex), $0.50 per task, 4.90 s TTFT. Against it,
164
- Astra low costs 64% more per task and scores 7 index points higher, with no
165
- published Astra coding-agent score at that level.
166
-
167
- Two risks come with it. Astra low loses AA-Briefcase, the eval closest to this
168
- kit's agent workload, and AA publishes no Coding Agent Index score for any Astra
169
- level except max. Watch reviewer output, because `reviewer` and
170
- `security-reviewer` inherit `@task`. If review quality drops, pin those agents
171
- to `anthropic/claude-opus-5:high` rather than reverting the whole role.
172
-
173
- ## What resolves to Astra low in practice
174
-
175
- `modelRoles.task` sets the model. The `:low` suffix alone does not cap every
176
- spawn. `SoT/.omp/models.yml` restricts the `openai-codex/gpt-6-astra` effort
177
- ladder to `[low]` through
178
- `providers.openai-codex.modelOverrides.gpt-6-astra.thinking`. omp clamps any
179
- requested effort to the model ladder, and `auto` has only one choice. Every
180
- Astra spawn therefore runs `low`, regardless of its `effort` hint.
171
+ Astra low beat Sol high on intelligence, latency, and token use at effectively
172
+ equal cost per task. It cost $0.82 against $0.81 and used 4k output tokens
173
+ per task against 13k. Its 2.60 s TTFT beat Sol high's 11.26 s.
174
+ That was a latency and token-budget win, not a cost saving.
175
+ The owner reversed that trade on purpose to reserve Astra for interactive
176
+ orchestration.
177
+
178
+ Sol high scores 64 in the Codex Coding Agent Index and 1361 on AA-Briefcase.
179
+ Astra low scores 1253 on AA-Briefcase. AA publishes no Astra Coding Agent
180
+ Index except max, which scores 67 in Codex.
181
+ The new Astra xhigh cycle stop scores 53 on the Intelligence Index, costs
182
+ $2.31 per index task, and has 161.65 s TTFT.
183
+ It has no published Coding Agent Index.
184
+
185
+ ## How Astra is selected in practice
186
+
187
+ No subagent role or `task.agentModelOverrides` entry resolves Astra.
188
+ `modelRoles.astra` is reachable only through the model switcher as a main
189
+ orchestrator role. The explicit Fable retry chain can also select Astra.
190
+ No subagent selects Astra automatically.
191
+
192
+ `SoT/.omp/models.yml` declares the full `low, medium, high, xhigh, max` ladder
193
+ with `defaultLevel: xhigh` under
194
+ `providers.openai-codex.modelOverrides.gpt-6-astra.thinking`.
195
+ The override stays as the kit's worked example of a provider ladder
196
+ redefinition. The in-session thinking control can still select `low` for a
197
+ quick answer.
181
198
 
182
199
  - The bundled `scout` and `sonic` agents carry `model: "@smol"` and
183
200
  `thinking-level: medium` in their embedded frontmatter, so they run Luna,
184
- not Astra. To move them, change `modelRoles.smol` or add a
201
+ not Sol or Astra. To move them, change `modelRoles.smol` or add a
185
202
  `task.agentModelOverrides` entry for the agent name.
186
203
  - The bundled `task` agent carries `model: "@task"` and
187
- `thinking-level: auto`. `auto` classifies each prompt and picks a level for
188
- the resolved model, but Astra's ladder permits only `low`.
204
+ `thinking-level: auto`. It resolves Sol, and `auto` classifies each prompt
205
+ to choose a thinking level.
189
206
  - `task.enableEffort` is `true`, so a caller can pass `effort: lo`, `med`, or
190
207
  `hi`, which overrides `auto`.
191
208
  - `task.maxEffort` is `high`, so `scout` and `sonic` run Luna `medium` by
192
209
  default and Luna `high` with `effort: hi`.
210
+ - The bundled `reviewer` and `security-reviewer` inherit `@task`, now Sol high.
193
211
  - The `code-reviewer` and `plan-reviewer` override entries remain dormant.
194
- omp's task tool rejects both names as unknown agents, so neither can spawn.
195
-
196
- Fresh `omp -p` runs on 2026-09-09 verified the model ladder with
197
- `task.maxEffort: high`. Child session logs recorded these results:
212
+ Both point to `@task`, now Sol high. omp's task tool rejects both names as
213
+ unknown agents, so neither can spawn.
198
214
 
199
- | Agent | Effort hint | Resolved model and level |
200
- |---|---|---|
201
- | `task` | `hi` | `gpt-6-astra:low` |
202
- | `task` | None, complex prompt | `gpt-6-astra:low` |
203
- | `reviewer` | `hi` | `gpt-6-astra:low` |
204
- | `security-reviewer` | None | `gpt-6-astra:low` |
205
- | `scout` | `hi` | `gpt-5.6-luna:high` |
206
- | `sonic` | `hi` | `gpt-5.6-luna:high` |
207
-
208
- The complex `task` run completed with 682 output tokens.
215
+ The runtime per-agent measurement from fresh `omp -p` runs on 2026-09-09
216
+ predates this change. It applies to the retired Astra-low `task` configuration,
217
+ not the current role map.
209
218
 
210
219
  omp's 272k context window is the `/extended-context off` window for Astra.
211
220
  `/extended-context on` uses the 922k input window, 1.05M total, which matches
@@ -214,8 +223,9 @@ AA's 1M.
214
223
  ## Free session launcher (`docks-kit omp`)
215
224
 
216
225
  `docks-kit omp [--model <selector>|--pick] [args...]` starts one interactive
217
- omp session on a single free model. All 12 model roles resolve to that model.
218
- All 9 retry fallback chains are empty, so a retry cannot reach a paid model.
226
+ omp session with a configuration overlay that selects a free model. The
227
+ overlay sets all 13 model roles to that model and empties all 10 retry fallback
228
+ chains. Higher-precedence model selection can replace those values.
219
229
  The overlay also sets `defaultThinkingLevel` and `task.maxEffort`, except for
220
230
  a model that publishes no thinking ladder, where both keys are omitted and the
221
231
  deployed values apply.
@@ -235,6 +245,28 @@ Every argument after the launcher flags forwards verbatim to omp. `docks-kit
235
245
  omp -p "..."` runs one prompt. `docks-kit omp --continue` resumes the
236
246
  previous session. A bare `docks-kit omp` opens an interactive session.
237
247
 
248
+ The overlay ranks above the deployed global and project configuration, and
249
+ below runtime overrides and later overlay files. Each input below therefore
250
+ selects a model that the overlay does not control:
251
+
252
+ - A forwarded later config overlay. The launcher passes its own `--config`
253
+ first, and omp applies a later overlay file over an earlier one.
254
+ - The runtime model flags `--model`, `--smol`, `--slow`, and `--plan`. The
255
+ legacy `--provider` flag selects a provider. `--models` sets the model
256
+ patterns that `Ctrl+P` cycling can reach, so cycling can leave the free
257
+ model.
258
+ - The model environment variables `PI_SMOL_MODEL`, `PI_SLOW_MODEL`, and
259
+ `PI_PLAN_MODEL`. `PI_CONFIG_FILES` is not in this set, because those files
260
+ load before `--config` overlays.
261
+ - The in-session model selector. `Alt+M` sets the roles, and `Alt+P` picks a
262
+ model for the current session only.
263
+ - A restored persisted session model. `--continue`, `--resume`, and
264
+ `autoResume` restore the model of the session they open, so a session that
265
+ started under the paid configuration returns to its paid model.
266
+
267
+ This list names the known selection paths. It is not a closed set. The
268
+ launcher does not reject these inputs today.
269
+
238
270
  The default model is `opencode-zen/muse-spark-1.3-contributor-free` ("Muse
239
271
  Spark 1.3 Free", 1,048,576-token context) at `xhigh`. `xhigh` is the ceiling
240
272
  of that model ladder: `minimal, low, medium, high, xhigh`. The paid
@@ -254,21 +286,44 @@ fail when the session starts.
254
286
 
255
287
  The choice persists per machine in `~/.docks-kit/state.json` under
256
288
  `ompSession`, next to `harnesses`. It survives across sessions. It is never
257
- committed. `--model <selector>` records one selector. `--pick` opens an
258
- interactive picker.
289
+ committed. `--model <selector>` records one selector and derives both levels
290
+ from the catalog row. `--pick` opens an interactive wizard.
259
291
 
260
292
  The picker lists only models the live `omp models --json` catalog reports at
261
- zero input and output cost (26 entries on 2026-09-11). This path cannot start
262
- a paid session. The catalog can advertise a free model that the account
263
- cannot call; omp reports that provider error unchanged.
293
+ zero input and output cost (26 entries on 2026-09-11). The resulting overlay
294
+ sets every role to the chosen free model and empties every retry chain.
295
+ Higher-precedence model selection can replace those values. The catalog can
296
+ advertise a free model that the account cannot call; omp reports that provider
297
+ error unchanged.
298
+
299
+ After the model, the wizard asks about thinking levels. `ompOverlay.ts
300
+ planEffortChoice` decides which questions apply, because omp accepts a
301
+ `:level` suffix only for a level the chosen model publishes:
302
+
303
+ - No ladder: the wizard asks nothing and the overlay omits every level.
304
+ - One level: the wizard states that level and asks nothing.
305
+ - Two or more levels: the wizard asks whether every role uses the same
306
+ level. Yes takes the highest level the model offers. No asks one level for
307
+ all roles except the advisor, then one level for the advisor.
308
+
309
+ The advisor question leads with a recommendation from `ompOverlay.ts
310
+ advisorRecommendation`: two steps down the model own ladder, clamped to the
311
+ lowest level that model publishes. The recommendation is the first row
312
+ because `Prompt.Select` starts on the first entry and accepts no initial
313
+ index, so `Enter` takes it. The full ladder follows in ladder order. A
314
+ recommendation equal to the chosen level is omitted rather than duplicated.
315
+
316
+ A session whose advisor differs from the other roles reports both on the
317
+ launch line.
264
318
 
265
319
  An omp login is required. The kit owns no login flow. It surfaces omp's own
266
320
  authentication error unchanged.
267
321
 
268
322
  When the catalog renames the free variant, change only the default selector
269
- constant. A ladder change needs no code change, because both levels come from
270
- the catalog row of the chosen model. Refresh the recorded default and the
271
- ladder counts in these docs in the same commit.
323
+ constant. A ladder change needs no code change. Under `--model` both levels
324
+ are derived from the catalog row of the chosen model, and under `--pick` the
325
+ user chooses them from the ladder that same row publishes. Refresh the
326
+ recorded default and the ladder counts in these docs in the same commit.
272
327
 
273
328
  ## Maintenance
274
329
 
@@ -278,5 +333,6 @@ ladder counts in these docs in the same commit.
278
333
  - Record the index version with the numbers. AA changes index composition
279
334
  between versions, so a score from another version is not a comparison.
280
335
  - Update the capture date in the same commit as any number.
281
- - Verify changes to the Astra ladder in `SoT/.omp/models.yml` and to
282
- `task.maxEffort` together with a fresh `omp -p` spawn per bundled agent.
336
+ - Verify Astra ladder changes in `SoT/.omp/models.yml` through the model
337
+ switcher and in-session thinking control.
338
+ - Verify `task.maxEffort` changes with a fresh `omp -p` spawn per bundled agent.