docks-kit 0.17.1 → 0.17.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +36 -32
- package/cli/docs/omp-models.md +88 -53
- package/cli/src/argv.ts +38 -469
- package/cli/src/argvSurface.ts +371 -0
- package/cli/src/argvValidate.ts +122 -0
- package/cli/src/commands/docs.ts +57 -42
- package/cli/src/commands/harnesses.ts +38 -45
- package/cli/src/commands/model.ts +48 -53
- package/cli/src/commands/models.ts +24 -26
- package/cli/src/commands/omp.ts +132 -112
- package/cli/src/commands/plugins.ts +20 -20
- package/cli/src/commands/skills.ts +22 -20
- package/cli/src/commands/status.ts +87 -80
- package/cli/src/commands/sync.ts +113 -87
- package/cli/src/commands/toolchain.ts +21 -21
- package/cli/src/commands/update.ts +139 -117
- package/cli/src/efforts.ts +35 -34
- package/cli/src/engine-native/DESIGN.md +30 -30
- package/cli/src/engine-native/bun.ts +77 -55
- package/cli/src/engine-native/claudeLsp.ts +154 -0
- package/cli/src/engine-native/claudeOptionalPlugins.ts +113 -0
- package/cli/src/engine-native/claudePluginPasses.ts +348 -0
- package/cli/src/engine-native/claudeRemovals.ts +195 -0
- package/cli/src/engine-native/claudeRetired.ts +4 -4
- package/cli/src/engine-native/claudeRuntime.ts +94 -68
- package/cli/src/engine-native/claudeSettings.ts +174 -0
- package/cli/src/engine-native/claudeSettingsModifiers.ts +50 -54
- package/cli/src/engine-native/claudeSync.ts +176 -456
- package/cli/src/engine-native/codexConfig.ts +158 -0
- package/cli/src/engine-native/codexHooks.ts +91 -0
- package/cli/src/engine-native/codexPlugins.ts +290 -0
- package/cli/src/engine-native/codexStatus.ts +24 -0
- package/cli/src/engine-native/codexSync.ts +110 -568
- package/cli/src/engine-native/codexToml.ts +140 -138
- package/cli/src/engine-native/deps.ts +144 -107
- package/cli/src/engine-native/engineCtx.ts +126 -0
- package/cli/src/engine-native/exec.ts +98 -87
- package/cli/src/engine-native/failures.ts +3 -3
- package/cli/src/engine-native/harnesses.ts +62 -72
- package/cli/src/engine-native/index.ts +34 -315
- package/cli/src/engine-native/jq.ts +25 -24
- package/cli/src/engine-native/logger.ts +98 -100
- package/cli/src/engine-native/models.ts +50 -45
- package/cli/src/engine-native/modes.ts +83 -80
- package/cli/src/engine-native/ompFileDeploy.ts +111 -0
- package/cli/src/engine-native/ompMarketplace.ts +136 -0
- package/cli/src/engine-native/ompOverlay.ts +67 -64
- package/cli/src/engine-native/ompPaths.ts +49 -38
- package/cli/src/engine-native/ompPlugins.ts +170 -0
- package/cli/src/engine-native/ompSync.ts +64 -408
- package/cli/src/engine-native/ompYaml.ts +43 -41
- package/cli/src/engine-native/os/darwin.ts +5 -5
- package/cli/src/engine-native/os/index.ts +15 -15
- package/cli/src/engine-native/os/linux.ts +5 -5
- package/cli/src/engine-native/os/posix.ts +28 -24
- package/cli/src/engine-native/os/targets.ts +17 -17
- package/cli/src/engine-native/os/types.ts +31 -31
- package/cli/src/engine-native/os/windows.ts +61 -61
- package/cli/src/engine-native/parseArgs.ts +122 -372
- package/cli/src/engine-native/parseHelp.ts +66 -0
- package/cli/src/engine-native/parseModifiers.ts +222 -0
- package/cli/src/engine-native/services.ts +34 -34
- package/cli/src/engine-native/settings.ts +15 -15
- package/cli/src/engine-native/sharedTypes.d.ts +46 -0
- package/cli/src/engine-native/skillsInstall.ts +105 -0
- package/cli/src/engine-native/skillsLinks.ts +203 -0
- package/cli/src/engine-native/skillsManifest.ts +16 -0
- package/cli/src/engine-native/skillsPrune.ts +108 -0
- package/cli/src/engine-native/skillsSync.ts +28 -362
- package/cli/src/engine-native/syncConcurrency.ts +55 -0
- package/cli/src/engine-native/syncDispatch.ts +141 -0
- package/cli/src/engine-native/toolchain.ts +65 -62
- package/cli/src/engine.ts +65 -57
- package/cli/src/generated/sotPayload.ts +7 -7
- package/cli/src/kitHome.ts +44 -40
- package/cli/src/main.ts +63 -56
- package/cli/src/manifests.ts +53 -54
- package/cli/src/md.d.ts +2 -2
- package/cli/src/payload.ts +9 -9
- package/cli/src/services.ts +21 -13
- package/cli/tsconfig.json +4 -1
- package/package.json +9 -2
- package/cli/src/engine-native/claudePlugins.ts +0 -510
package/AGENTS.md
CHANGED
|
@@ -33,6 +33,10 @@ launcher can fall back to Bun source.
|
|
|
33
33
|
| `cli/src/engine-native/ompSync.ts` | omp file deployment, marketplace registration, and plugin synchronization |
|
|
34
34
|
| `cli/src/commands/omp.ts` | `docks-kit omp` session launcher: renders the free-model run overlay and forwards args to omp |
|
|
35
35
|
| `cli/src/engine-native/ompOverlay.ts` | Free-model overlay render plus catalog parse and thinking-ceiling helpers |
|
|
36
|
+
| `cli/src/engine-native/<axis>.ts` | One change axis per file: `codexConfig`/`codexHooks`/`codexPlugins`/`codexStatus`, `claudeSettings`/`claudeRemovals`/`claudePluginPasses`/`claudeOptionalPlugins`/`claudeLsp`, `parseModifiers`/`parseHelp`, `ompFileDeploy`/`ompMarketplace`/`ompPlugins`, `skillsManifest`/`skillsLinks`/`skillsInstall`/`skillsPrune`, `engineCtx`/`syncConcurrency`/`syncDispatch`; orchestrators keep prior exports |
|
|
37
|
+
| `cli/src/engine-native/sharedTypes.d.ts` | Single definition per shared shape (manifest records, settings edits, omp session model, scalar modifier flags) |
|
|
38
|
+
| `.oxlintrc.json` / `.oxfmtrc.json` | Mechanical lint (correctness, suspicious, eqeqeq, no-var, prefer-const, no-unused) and code format scope |
|
|
39
|
+
| `cli/test/fixtures/` | Golden fixtures: `home-fresh`, `home-drift`, `home-invalid-json`, `codex-toml`, `statusline` |
|
|
36
40
|
| `cli/` | Effect 4 RC CLI + bundled docs topics |
|
|
37
41
|
| `SoT/models.json` | Kit-verified Claude and Codex model catalog |
|
|
38
42
|
| `SoT/toolchain.json` | Toolchain floors manifest (verified pins consumed by EngineNative) |
|
|
@@ -66,15 +70,15 @@ omp SoT notes:
|
|
|
66
70
|
- `SoT/.omp/AGENTS.md`, `config.yml`, `models.yml`, and `mcp.json` deploy to `~/.omp/agent/`.
|
|
67
71
|
- `SoT/.omp/intercom.json` deploys to `$PI_CODING_AGENT_DIR/intercom/config.json`. The default root is `~/.pi/agent`.
|
|
68
72
|
- `ompSync.ts syncMergedYaml` deep-merges `config.yml` through `ompYaml.ts mergeOmpConfig` and `models.yml` through `mergeOmpModels`. Both wrap one generic mapping merge; only the config wrapper prunes stale `retry.fallbackChains` wildcards.
|
|
69
|
-
- `cycleOrder` ends with
|
|
70
|
-
- `modelRoles.task` is `openai-codex/gpt-6-
|
|
71
|
-
- `SoT/.omp/models.yml`
|
|
73
|
+
- `cycleOrder` ends with `astra` as its fifth stop. `modelRoles.astra` is `openai-codex/gpt-6-astra:xhigh`, and `modelTags.astra` is visible. Astra and Fable fall back to each other through concrete selectors. `modelRoles.fable` remains `anthropic/claude-fable-5-1:medium`, with visible `modelTags.fable`. The hidden `switch_fable` role uses the same Fable selector and keeps an empty fallback chain.
|
|
74
|
+
- `modelRoles.task` is `openai-codex/gpt-5.6-sol:high`. The four reviewer entries in `task.agentModelOverrides` inherit it through `@task`. Only bundled `reviewer` and `security-reviewer` are discoverable OMP agents; `code-reviewer` and `plan-reviewer` stay dormant.
|
|
75
|
+
- `SoT/.omp/models.yml` declares Astra's full `low, medium, high, xhigh, max` ladder with `defaultLevel: xhigh` as the worked provider ladder-override example. `ompYaml.ts mergeOmpModels` preserves deployed-only keys in `~/.omp/agent/models.yml`, because a user file may carry provider credentials. Whole-file replacement is wrong.
|
|
72
76
|
- `SoT/.omp/AGENTS.md` carries the rule `Please remove all mannered prose.` Anthropic's Fable 5.1 prompting guide documents mannered prose as a Fable 5.1 behavior and gives that sentence as its short-version fix: https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1
|
|
73
77
|
- `cli/docs/omp-models.md` (topic `omp-models`) records the role map rationale and the Artificial Analysis snapshot behind it. Model choices change with published benchmarks, so update that topic in the same commit as a role change.
|
|
74
78
|
- Sync registers the `docks` marketplace. It installs or upgrades `docks@docks` and `plan-lifecycle@docks` at user scope.
|
|
75
79
|
- Sync installs `pi-intercom` at the verified version from `SoT/toolchain.json`.
|
|
76
80
|
- The omp CLI is upstream-owned and self-updating through `omp update`. Sync never installs or upgrades the CLI.
|
|
77
|
-
- `docks-kit omp [--model <selector>|--pick] [args...]` starts one interactive omp session
|
|
81
|
+
- `docks-kit omp [--model <selector>|--pick] [args...]` starts one interactive omp session with every model role set to one free model and every retry chain empty. Higher-precedence model selection can replace those overlay values. It renders the overlay to `~/.cache/docks-kit/omp-free-<model>-<digest>.yml` (mode 0600) and passes it through omp's repeatable `--config` flag. `ompOverlay.ts overlayFileName` gives each model its own file, because omp can re-read the overlay during a live session and a second launcher on another model must not rewrite it. Deployed `~/.omp/agent/` files stay untouched, so the next plain `omp` run uses the paid configuration again. Remaining arguments forward verbatim to omp.
|
|
78
82
|
- The session model persists per machine in `~/.docks-kit/state.json` under `ompSession`, next to `harnesses`. The default is `opencode-zen/muse-spark-1.3-contributor-free` at `xhigh`. `ompOverlay.ts ladderCeiling, advisorLevelFor` derive both levels from the ladder of the chosen model for the `--model` path, and `ompOverlay.ts planEffortChoice` drives the `--pick` wizard, which asks whether every role shares one level and otherwise takes one level for the main roles and one for the advisor from that model's own ladder. Levels never come from a fixed list: free ladders are not uniform, three free models publish no `medium`, and two publish no ladder, which the overlay renders as bare selectors. The picker lists only zero-cost catalog models.
|
|
79
83
|
|
|
80
84
|
For per-tool SoT layouts (`SoT/.claude/`, `SoT/.codex/`, `SoT/.omp/`), see the matching SoT directory.
|
|
@@ -90,8 +94,10 @@ For per-tool SoT layouts (`SoT/.claude/`, `SoT/.codex/`, `SoT/.omp/`), see the m
|
|
|
90
94
|
- **Per-machine harness selection.** `~/.docks-kit/state.json` drives a flag-less sync. A missing file selects `claude`, `codex`, and `agents`; it never selects `omp` implicitly. `sync` never prompts and never writes the selection file. `docks-kit harnesses` is the only command that writes the `harnesses` selection. The same state file also carries the independent `ompSession` key, which only `docks-kit omp` writes; every writer merges over the stored record, so neither key can drop the other.
|
|
91
95
|
- **Additive by default.** Keys present in deployed config but absent from SoT are preserved on default sync. This protects user-only additions, but means drift accumulates — neither flag-less reset can clean it up. The one exception is the Claude `removed` manifest (`claudeSync.ts REMOVED_MANIFEST, baseline removal inventory`), a curated list of unambiguous kit-owned artifacts that `claudeSync.ts syncRemovals, baseline artifact prune` force-prunes on every sync, including the home-relative `~/.local/bin/session-relay` artifact installed outside `~/.claude`; see `CLAUDE.md` § Pruning stale artifacts.
|
|
92
96
|
- **`--reconcile` / `--prune` are the kit-owned reconcile flags.** Orthogonal — `--reconcile` reconciles the settings layer (SoT-declared keys/tables/arrays win; user-only keys and nested objects are preserved; permissions arrays are replaced wholesale by SoT). `--prune` uninstalls kit-managed installations not in the SoT (plugins, marketplaces, and `~/.agents/skills/*` entries tracked in `~/.agents/.kit-managed-skills`). Combine for a full reset to SoT's kit-managed scope. User-only additions outside the kit's scope (custom env vars, mcpServers, manually-installed skills, third-party plugins not declared in SoT) are always preserved. Each tool's per-tool file documents the specific paths and diff recipes.
|
|
93
|
-
- **SOLID-aligned modules.**
|
|
97
|
+
- **SOLID-aligned modules.** Axis modules own one concern each (see layout); orchestrators (`claudeSync.ts`, `codexSync.ts`, `skillsSync.ts`, `ompSync.ts`, `index.ts`, `argv.ts`, `parseArgs.ts`) keep prior exports. New shared shapes go in `sharedTypes.d.ts`, never a second local copy. The public CLI seam is `cli/src/engine.ts`.
|
|
94
98
|
- **Small, reviewable changes.** Bundled multi-concern PRs are harder to review and revert. Split an engine/CLI change and a per-tool config change unless the change requires atomicity.
|
|
99
|
+
- **Version bumps.** Bump `package.json` on user-visible behavior change; bump a `SoT/toolchain.json` `verified` pin only after testing that release. Refactors, splits, and doc edits bump nothing.
|
|
100
|
+
- **Regenerate the payload.** After `SoT/`, `notification.mp3`, `package.json`, or `cli/tsconfig.json` changes, run `bun cli/scripts/generate-sot-payload.ts`; `bun run check:generated` verifies it. After file moves, clamp cited `metadata.source_files[].lines` ranges so `skillMetadata.test.ts` passes.
|
|
95
101
|
- **Dry-run before destructive flags.** Always preview with `./docks-kit sync --dry-run` (or the relevant `diff <(jq -S …)` recipe in the per-tool file) before invoking `--reconcile` or `--prune`. User-added permissions / env vars / plugins absent from SoT will be discarded.
|
|
96
102
|
- **SoT prompt files are rules, not explanation.** `SoT/.claude/CLAUDE.md`, `SoT/.codex/AGENTS.md`, and `SoT/.omp/AGENTS.md` are loaded into agent sessions' prompt context — every line costs prompt tokens on every turn for every user. Restrict their content to rules, heuristics, and `<constraint>` blocks the agent must *act on* during a turn. Do NOT add inline source citations (`Source: …`, attributed quotes), "why this rule exists" preface text, version-watermarking trivia (e.g. "Distilled from X v2.0, captured 2025-11-07"), per-bug workarounds, or installation instructions. Provenance, motivation, and historical context belong in `CLAUDE.md` / `AGENTS.md` at the repo root (humans read once) or in commit messages — never in the SoT. For every line, apply the official test: …
|
|
97
103
|
- **Cache-invariance for kit-authored prompt surfaces.** Never put timestamps, counters, or mutable state into SoT prompt files, hook outputs that land in the cached prefix, or tool definitions — cache breaks force cold-start writes. Dynamic context belongs in runtime-injected messages (e.g. SessionStart hook output), which is exactly how the kit's date/config injection works.
|
|
@@ -109,7 +115,7 @@ For per-tool SoT layouts (`SoT/.claude/`, `SoT/.codex/`, `SoT/.omp/`), see the m
|
|
|
109
115
|
- **One exemption: the kit's own package.** `install.sh` and `install.ps1` end with `bun add -g docks-kit@latest`, because a global installer that pinned itself would install a fixed old kit forever, and pinning it to `package.json` would request an unpublished version between the release-prep commit and the npm publish. The exemption covers `docks-kit` alone. Both installers still pin the Bun installer they download to the manifest's verified version, and `cli/test/unit/install.test.ts` asserts that pin in all four launcher and installer scripts.
|
|
110
116
|
## Testing
|
|
111
117
|
|
|
112
|
-
|
|
118
|
+
Run `bun run check` before and after each change: oxlint, `tsc --noEmit`, unit tests, and both golden suites. Scripts: `lint`, `typecheck`, `test:unit`, `test:integration` (goldens), `test`, `format`; CI runs `check` on push and pull request with Bun 1.4.2. Prove-red modes must exit non-zero after detecting planted mismatches. Also verify user-facing changes via `./docks-kit sync --dry-run`, per-tool sanity (`/doctor`, `/plugin`, etc.), `bun run test:runtime:posix`, and `diff <(jq -S . <SoT>) <(jq -S . <deployed>)` recipes from the per-tool file.
|
|
113
119
|
|
|
114
120
|
Use direct acceptance and focused regressions while iterating, then run the full unit/golden gate once at the pre-commit or release boundary. Reuse still-matching evidence; a later relevant edit invalidates only the affected rung and final gate, not every prior check.
|
|
115
121
|
|
|
@@ -142,7 +148,7 @@ The Output Standard wins.
|
|
|
142
148
|
Use `unslop` only for its pattern-detection lists.
|
|
143
149
|
|
|
144
150
|
<constraint>
|
|
145
|
-
When a kit-mechanic skill, its `references/`, or a wrapper agent (`.claude/agents/*.md` + its `.codex/agents/*.toml` twin) cites EngineNative internals, name the **module + exported/local function + semantic anchor** (e.g. `
|
|
151
|
+
When a kit-mechanic skill, its `references/`, or a wrapper agent (`.claude/agents/*.md` + its `.codex/agents/*.toml` twin) cites EngineNative internals, name the **module + exported/local function + semantic anchor** (e.g. `claudePluginPasses.ts syncPlugins, pass 5 uninstall guard`) — never a raw `file:NNN` line number, which goes stale on every refactor. Keep exactly one coarse `metadata.source_files[].lines` range per skill file as the sole intentional line-number touchpoint.
|
|
146
152
|
</constraint>
|
|
147
153
|
|
|
148
154
|
**Universal-skill bootstrap.** `SoT/.agents/skills.txt` is intentionally empty, so the default sync exposes no universal skills. The generic opt-in contract remains: each [agentskills.io](https://agentskills.io/specification) slug is installed by `skillsSync.ts` with `npx skills add <slug> -g -y -a claude-code codex`. The slug comes first because `-a/--agent` is variadic; naming both supported agents preserves the canonical `~/.agents/skills/<name>/SKILL.md` plus Claude's `~/.claude/skills/<name>` symlink. Existing canonical directories are reused idempotently, `--prune` removes only entries tracked in `~/.agents/.kit-managed-skills`, and skills requiring a separate CLI retain explicit helpers in `skillsSync.ts`.
|
|
@@ -158,33 +164,31 @@ non-`local` effect.
|
|
|
158
164
|
|
|
159
165
|
<constraint>
|
|
160
166
|
The plan record is a GitHub issue. Its body starts with
|
|
161
|
-
`<!-- plan-contract:
|
|
162
|
-
sections
|
|
163
|
-
timestamps, and completion
|
|
164
|
-
Exactly three skills own the workflow
|
|
165
|
-
|
|
166
|
-
review, implement, code review
|
|
167
|
-
review phases, then archives
|
|
168
|
-
pre-implementation verdict block per round. Two read-only reviewer
|
|
169
|
-
ship
|
|
170
|
-
a wrapper.
|
|
167
|
+
`<!-- plan-contract: v4 -->`, followed by one blank line and seven `##`
|
|
168
|
+
sections. It has no frontmatter. GitHub owns the title, open-work phase, owner,
|
|
169
|
+
timestamps, and completion. The repository tracks no plan markdown.
|
|
170
|
+
Exactly three skills own the workflow. `plan-workspace` maintains the workspace.
|
|
171
|
+
Main-context `plan-manager` runs six phases: decide, draft, research, plan
|
|
172
|
+
review, implement, and code review. It uses bounded repair and fresh re-review
|
|
173
|
+
in both review phases, then archives. Internal `plan-reviewer` returns one
|
|
174
|
+
readable pre-implementation verdict block per round. Two read-only reviewer
|
|
175
|
+
wrappers ship: `plan-reviewer` and `code-reviewer`.
|
|
171
176
|
</constraint>
|
|
172
177
|
|
|
173
|
-
After the marker and blank line, the record carries exactly
|
|
174
|
-
|
|
175
|
-
`##
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
|
|
179
|
-
|
|
180
|
-
|
|
181
|
-
|
|
182
|
-
|
|
183
|
-
must equal the plan's sole assignee.
|
|
184
|
-
|
|
185
|
-
|
|
186
|
-
|
|
187
|
-
plan-review `blocked` verdict always routes its user-only decision through
|
|
178
|
+
After the marker and blank line, the record carries exactly seven sections.
|
|
179
|
+
They appear once each in this order: `## Goal`, `## Research`, `## Steps`,
|
|
180
|
+
`## Acceptance`, `## Do not touch`, `## Open questions`, and
|
|
181
|
+
`## Verification Results`. `## Goal` carries exactly one mode line.
|
|
182
|
+
Open-work phase is `drafting`, `planned`, `ongoing`, or `blocked`.
|
|
183
|
+
Every plan carries `plan`. A `plan:<phase>` label stores the open phase.
|
|
184
|
+
A blocked plan starts `## Open questions` with `Blocked: <one-line reason>`.
|
|
185
|
+
GitHub `state` and `stateReason` determine closed status. Review records live
|
|
186
|
+
only in issue comments. Each reviewer returns one markdown block.
|
|
187
|
+
The manager posts that block as one issue comment. The latest trusted eligible
|
|
188
|
+
comment per review kind wins. Its author must equal the plan's sole assignee.
|
|
189
|
+
Both review phases use fresh inputs and run at most five rounds. They stop on
|
|
190
|
+
pass, no progress, a finding that survives repair, or a round-five non-pass.
|
|
191
|
+
A plan-review `blocked` verdict routes its user-only decision through
|
|
188
192
|
`## Open questions` and `ask`.
|
|
189
193
|
|
|
190
194
|
The record carries no hash, permit, run identity, lock, or bundle, and the
|
package/cli/docs/omp-models.md
CHANGED
|
@@ -11,7 +11,7 @@ against.
|
|
|
11
11
|
| `default` | `anthropic/claude-opus-5` | high | 48 | $3.61 | 16.96 s | 66 (Claude Code) |
|
|
12
12
|
| `slow` | `anthropic/claude-opus-5` | xhigh | 50 | $4.88 | 28.65 s | 68 (Claude Code) |
|
|
13
13
|
| `plan` | `anthropic/claude-opus-5` | xhigh | 50 | $4.88 | 28.65 s | 68 (Claude Code) |
|
|
14
|
-
| `task` | `openai-codex/gpt-6-
|
|
14
|
+
| `task` | `openai-codex/gpt-5.6-sol` | high | 42 | $0.81 | 11.26 s | 64 (Codex) |
|
|
15
15
|
| `advisor` | `openai-codex/gpt-5.6-sol` | medium | 39 | $0.50 | 4.90 s | 62 (Codex) |
|
|
16
16
|
| `designer` | `anthropic/claude-opus-5` | high | 48 | $3.61 | 16.96 s | 66 (Claude Code) |
|
|
17
17
|
| `vision` | `anthropic/claude-opus-5` | medium | 45 | $2.19 | 3.79 s | 64 (Claude Code) |
|
|
@@ -19,6 +19,7 @@ against.
|
|
|
19
19
|
| `tiny` | `openai-codex/gpt-5.6-luna` | low | 22 | n/a | 1.78 s | 25 (Codex) |
|
|
20
20
|
| `fable` | `anthropic/claude-fable-5-1` | medium | 49 | $2.98 | 9.65 s | n/a |
|
|
21
21
|
| `switch_fable` | `anthropic/claude-fable-5-1` | medium | 49 | $2.98 | 9.65 s | n/a |
|
|
22
|
+
| `astra` | `openai-codex/gpt-6-astra` | xhigh | 53 | $2.31 | 161.65 s | n/a |
|
|
22
23
|
|
|
23
24
|
The table reports the measured Artificial Analysis figures for each assigned
|
|
24
25
|
model and level. It states no motive that the config or omp's own
|
|
@@ -28,8 +29,8 @@ max.
|
|
|
28
29
|
|
|
29
30
|
What omp's settings catalog establishes about these roles:
|
|
30
31
|
|
|
31
|
-
- `cycleOrder` lists the roles the model switcher cycles
|
|
32
|
-
fourth `Ctrl+P` stop.
|
|
32
|
+
- `cycleOrder` lists the roles the model switcher cycles. `fable` is the
|
|
33
|
+
fourth `Ctrl+P` stop, and `astra` is the fifth and last stop.
|
|
33
34
|
- `tiny` overrides the model for lightweight background tasks: titles, memory,
|
|
34
35
|
auto-thinking, and unexpected-stop detection.
|
|
35
36
|
- `modelTags` carries role metadata and can introduce roles; `hidden: true`
|
|
@@ -44,6 +45,14 @@ are dormant until an OMP agent with that name exists.
|
|
|
44
45
|
`retry.fallbackChains.task` keeps `anthropic/claude-opus-5:high` as a
|
|
45
46
|
cross-vendor fallback.
|
|
46
47
|
|
|
48
|
+
`retry.fallbackChains.astra` holds `anthropic/claude-fable-5-1:medium`.
|
|
49
|
+
`retry.fallbackChains.fable` holds `openai-codex/gpt-6-astra:xhigh`.
|
|
50
|
+
Each deliberate cycle stop falls to the other vendor. Without these explicit
|
|
51
|
+
chains, `retry.fallbackChains.default` would send either stop to
|
|
52
|
+
`openai-codex/gpt-5.6-sol:high`.
|
|
53
|
+
Chain entries are concrete selectors, not role aliases, so this pair cannot
|
|
54
|
+
recurse. The hidden `switch_fable` chain stays empty.
|
|
55
|
+
|
|
47
56
|
## Artificial Analysis snapshot
|
|
48
57
|
|
|
49
58
|
Source: `https://artificialanalysis.ai`, read on 2026-09-08. Every score below
|
|
@@ -139,9 +148,14 @@ published numeric cache price. Context 1M.
|
|
|
139
148
|
Price: $0.20 in, $1.20 out, $0.02 cache read per 1M; no published cache-write
|
|
140
149
|
price. Context 1M. AA lists cost per task only for max and xhigh.
|
|
141
150
|
|
|
142
|
-
## Why `task`
|
|
151
|
+
## Why `task` returns to Sol high
|
|
152
|
+
|
|
153
|
+
`task` runs `openai-codex/gpt-5.6-sol:high`. The owner uses Astra only for
|
|
154
|
+
main orchestration, so Astra now has a dedicated `astra` cycle stop at `xhigh`.
|
|
155
|
+
The comparison below records the retired Astra-low choice beside the current
|
|
156
|
+
Sol-high choice and the other measured alternatives.
|
|
143
157
|
|
|
144
|
-
| Metric | Astra low | Sol high | Sol max | Opus 5 high |
|
|
158
|
+
| Metric | Astra low, retired task | Sol high, current task | Sol max | Opus 5 high |
|
|
145
159
|
|---|---:|---:|---:|---:|
|
|
146
160
|
| Intelligence Index | 46 | 42 | 47 | 48 |
|
|
147
161
|
| Cost per Index task | $0.82 | $0.81 | $1.99 | $3.61 |
|
|
@@ -154,58 +168,53 @@ price. Context 1M. AA lists cost per task only for max and xhigh.
|
|
|
154
168
|
| AA-Briefcase | 1253 | 1361 | 1475 | 1557 |
|
|
155
169
|
| AA-Omniscience | 41 | 20 | 22 | 34 |
|
|
156
170
|
|
|
157
|
-
Astra low
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
|
|
164
|
-
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
|
|
172
|
-
|
|
173
|
-
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
`
|
|
179
|
-
|
|
180
|
-
|
|
171
|
+
Astra low beat Sol high on intelligence, latency, and token use at effectively
|
|
172
|
+
equal cost per task. It cost $0.82 against $0.81 and used 4k output tokens
|
|
173
|
+
per task against 13k. Its 2.60 s TTFT beat Sol high's 11.26 s.
|
|
174
|
+
That was a latency and token-budget win, not a cost saving.
|
|
175
|
+
The owner reversed that trade on purpose to reserve Astra for interactive
|
|
176
|
+
orchestration.
|
|
177
|
+
|
|
178
|
+
Sol high scores 64 in the Codex Coding Agent Index and 1361 on AA-Briefcase.
|
|
179
|
+
Astra low scores 1253 on AA-Briefcase. AA publishes no Astra Coding Agent
|
|
180
|
+
Index except max, which scores 67 in Codex.
|
|
181
|
+
The new Astra xhigh cycle stop scores 53 on the Intelligence Index, costs
|
|
182
|
+
$2.31 per index task, and has 161.65 s TTFT.
|
|
183
|
+
It has no published Coding Agent Index.
|
|
184
|
+
|
|
185
|
+
## How Astra is selected in practice
|
|
186
|
+
|
|
187
|
+
No subagent role or `task.agentModelOverrides` entry resolves Astra.
|
|
188
|
+
`modelRoles.astra` is reachable only through the model switcher as a main
|
|
189
|
+
orchestrator role. The explicit Fable retry chain can also select Astra.
|
|
190
|
+
No subagent selects Astra automatically.
|
|
191
|
+
|
|
192
|
+
`SoT/.omp/models.yml` declares the full `low, medium, high, xhigh, max` ladder
|
|
193
|
+
with `defaultLevel: xhigh` under
|
|
194
|
+
`providers.openai-codex.modelOverrides.gpt-6-astra.thinking`.
|
|
195
|
+
The override stays as the kit's worked example of a provider ladder
|
|
196
|
+
redefinition. The in-session thinking control can still select `low` for a
|
|
197
|
+
quick answer.
|
|
181
198
|
|
|
182
199
|
- The bundled `scout` and `sonic` agents carry `model: "@smol"` and
|
|
183
200
|
`thinking-level: medium` in their embedded frontmatter, so they run Luna,
|
|
184
|
-
not Astra. To move them, change `modelRoles.smol` or add a
|
|
201
|
+
not Sol or Astra. To move them, change `modelRoles.smol` or add a
|
|
185
202
|
`task.agentModelOverrides` entry for the agent name.
|
|
186
203
|
- The bundled `task` agent carries `model: "@task"` and
|
|
187
|
-
`thinking-level: auto`. `auto` classifies each prompt
|
|
188
|
-
|
|
204
|
+
`thinking-level: auto`. It resolves Sol, and `auto` classifies each prompt
|
|
205
|
+
to choose a thinking level.
|
|
189
206
|
- `task.enableEffort` is `true`, so a caller can pass `effort: lo`, `med`, or
|
|
190
207
|
`hi`, which overrides `auto`.
|
|
191
208
|
- `task.maxEffort` is `high`, so `scout` and `sonic` run Luna `medium` by
|
|
192
209
|
default and Luna `high` with `effort: hi`.
|
|
210
|
+
- The bundled `reviewer` and `security-reviewer` inherit `@task`, now Sol high.
|
|
193
211
|
- The `code-reviewer` and `plan-reviewer` override entries remain dormant.
|
|
194
|
-
omp's task tool rejects both names as
|
|
195
|
-
|
|
196
|
-
Fresh `omp -p` runs on 2026-09-09 verified the model ladder with
|
|
197
|
-
`task.maxEffort: high`. Child session logs recorded these results:
|
|
212
|
+
Both point to `@task`, now Sol high. omp's task tool rejects both names as
|
|
213
|
+
unknown agents, so neither can spawn.
|
|
198
214
|
|
|
199
|
-
|
|
200
|
-
|
|
201
|
-
|
|
202
|
-
| `task` | None, complex prompt | `gpt-6-astra:low` |
|
|
203
|
-
| `reviewer` | `hi` | `gpt-6-astra:low` |
|
|
204
|
-
| `security-reviewer` | None | `gpt-6-astra:low` |
|
|
205
|
-
| `scout` | `hi` | `gpt-5.6-luna:high` |
|
|
206
|
-
| `sonic` | `hi` | `gpt-5.6-luna:high` |
|
|
207
|
-
|
|
208
|
-
The complex `task` run completed with 682 output tokens.
|
|
215
|
+
The runtime per-agent measurement from fresh `omp -p` runs on 2026-09-09
|
|
216
|
+
predates this change. It applies to the retired Astra-low `task` configuration,
|
|
217
|
+
not the current role map.
|
|
209
218
|
|
|
210
219
|
omp's 272k context window is the `/extended-context off` window for Astra.
|
|
211
220
|
`/extended-context on` uses the 922k input window, 1.05M total, which matches
|
|
@@ -214,8 +223,9 @@ AA's 1M.
|
|
|
214
223
|
## Free session launcher (`docks-kit omp`)
|
|
215
224
|
|
|
216
225
|
`docks-kit omp [--model <selector>|--pick] [args...]` starts one interactive
|
|
217
|
-
omp session
|
|
218
|
-
|
|
226
|
+
omp session with a configuration overlay that selects a free model. The
|
|
227
|
+
overlay sets all 13 model roles to that model and empties all 10 retry fallback
|
|
228
|
+
chains. Higher-precedence model selection can replace those values.
|
|
219
229
|
The overlay also sets `defaultThinkingLevel` and `task.maxEffort`, except for
|
|
220
230
|
a model that publishes no thinking ladder, where both keys are omitted and the
|
|
221
231
|
deployed values apply.
|
|
@@ -235,6 +245,28 @@ Every argument after the launcher flags forwards verbatim to omp. `docks-kit
|
|
|
235
245
|
omp -p "..."` runs one prompt. `docks-kit omp --continue` resumes the
|
|
236
246
|
previous session. A bare `docks-kit omp` opens an interactive session.
|
|
237
247
|
|
|
248
|
+
The overlay ranks above the deployed global and project configuration, and
|
|
249
|
+
below runtime overrides and later overlay files. Each input below therefore
|
|
250
|
+
selects a model that the overlay does not control:
|
|
251
|
+
|
|
252
|
+
- A forwarded later config overlay. The launcher passes its own `--config`
|
|
253
|
+
first, and omp applies a later overlay file over an earlier one.
|
|
254
|
+
- The runtime model flags `--model`, `--smol`, `--slow`, and `--plan`. The
|
|
255
|
+
legacy `--provider` flag selects a provider. `--models` sets the model
|
|
256
|
+
patterns that `Ctrl+P` cycling can reach, so cycling can leave the free
|
|
257
|
+
model.
|
|
258
|
+
- The model environment variables `PI_SMOL_MODEL`, `PI_SLOW_MODEL`, and
|
|
259
|
+
`PI_PLAN_MODEL`. `PI_CONFIG_FILES` is not in this set, because those files
|
|
260
|
+
load before `--config` overlays.
|
|
261
|
+
- The in-session model selector. `Alt+M` sets the roles, and `Alt+P` picks a
|
|
262
|
+
model for the current session only.
|
|
263
|
+
- A restored persisted session model. `--continue`, `--resume`, and
|
|
264
|
+
`autoResume` restore the model of the session they open, so a session that
|
|
265
|
+
started under the paid configuration returns to its paid model.
|
|
266
|
+
|
|
267
|
+
This list names the known selection paths. It is not a closed set. The
|
|
268
|
+
launcher does not reject these inputs today.
|
|
269
|
+
|
|
238
270
|
The default model is `opencode-zen/muse-spark-1.3-contributor-free` ("Muse
|
|
239
271
|
Spark 1.3 Free", 1,048,576-token context) at `xhigh`. `xhigh` is the ceiling
|
|
240
272
|
of that model ladder: `minimal, low, medium, high, xhigh`. The paid
|
|
@@ -258,9 +290,11 @@ committed. `--model <selector>` records one selector and derives both levels
|
|
|
258
290
|
from the catalog row. `--pick` opens an interactive wizard.
|
|
259
291
|
|
|
260
292
|
The picker lists only models the live `omp models --json` catalog reports at
|
|
261
|
-
zero input and output cost (26 entries on 2026-09-11).
|
|
262
|
-
|
|
263
|
-
|
|
293
|
+
zero input and output cost (26 entries on 2026-09-11). The resulting overlay
|
|
294
|
+
sets every role to the chosen free model and empties every retry chain.
|
|
295
|
+
Higher-precedence model selection can replace those values. The catalog can
|
|
296
|
+
advertise a free model that the account cannot call; omp reports that provider
|
|
297
|
+
error unchanged.
|
|
264
298
|
|
|
265
299
|
After the model, the wizard asks about thinking levels. `ompOverlay.ts
|
|
266
300
|
planEffortChoice` decides which questions apply, because omp accepts a
|
|
@@ -299,5 +333,6 @@ recorded default and the ladder counts in these docs in the same commit.
|
|
|
299
333
|
- Record the index version with the numbers. AA changes index composition
|
|
300
334
|
between versions, so a score from another version is not a comparison.
|
|
301
335
|
- Update the capture date in the same commit as any number.
|
|
302
|
-
- Verify
|
|
303
|
-
|
|
336
|
+
- Verify Astra ladder changes in `SoT/.omp/models.yml` through the model
|
|
337
|
+
switcher and in-session thinking control.
|
|
338
|
+
- Verify `task.maxEffort` changes with a fresh `omp -p` spawn per bundled agent.
|