omnilane 0.34.0 → 0.42.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (42) hide show
  1. package/.claude-plugin/marketplace.json +4 -4
  2. package/.claude-plugin/plugin.json +2 -2
  3. package/CHANGELOG.md +71 -1
  4. package/README.ja.md +63 -33
  5. package/README.ko.md +63 -32
  6. package/README.md +147 -86
  7. package/README.zh-CN.md +61 -30
  8. package/README.zh-TW.md +124 -75
  9. package/VERSION +1 -1
  10. package/config/aa-model-policy.json +3046 -0
  11. package/docs/aa-model-coverage-2026-09-05.json +29204 -0
  12. package/docs/completion-wakeup.md +126 -0
  13. package/docs/model-capabilities-2026-09.md +380 -0
  14. package/docs/native-executor.md +264 -0
  15. package/docs/release-notes-0.42.1.md +32 -0
  16. package/hooks/routing-instruction.md +101 -40
  17. package/package.json +8 -2
  18. package/plugin.json +2 -2
  19. package/routing.local.yaml.example +8 -3
  20. package/routing.yaml +16 -16
  21. package/scripts/completion-wakeup.py +390 -0
  22. package/scripts/configure.sh +4 -4
  23. package/scripts/dispatch.sh +323 -32
  24. package/scripts/doctor.sh +55 -1
  25. package/scripts/jobs.sh +64 -17
  26. package/scripts/lib/aa_policy.py +473 -0
  27. package/scripts/lib/aa_retry.py +77 -0
  28. package/scripts/lib/common.sh +106 -1
  29. package/scripts/lib/job-worker.sh +314 -20
  30. package/scripts/lib/live-protocol.sh +147 -2
  31. package/scripts/lib/native.py +507 -0
  32. package/scripts/lib/normalize-claude-stream.py +72 -0
  33. package/scripts/lib/prepare-agy-mode.py +374 -0
  34. package/scripts/release-audit.sh +103 -0
  35. package/scripts/runners/run-claude.sh +81 -47
  36. package/scripts/runners/run-codex-live.py +462 -0
  37. package/scripts/runners/run-codex.sh +62 -3
  38. package/scripts/runners/run-gemini.sh +85 -10
  39. package/scripts/runners/run-grok-live.py +426 -0
  40. package/scripts/runners/run-grok.sh +117 -6
  41. package/scripts/runners/run-vote.sh +6 -3
  42. package/skills/omnilane/SKILL.md +217 -81
@@ -7,23 +7,18 @@ description: 'Universal model-routing table + cross-vendor dispatch for ANY harn
7
7
 
8
8
  You (the main loop) may be Claude, GPT, Grok, or Gemini. The procedure is identical:
9
9
 
10
- 1. **Identify your main model.** You know which model you are running as.
10
+ 1. **Identify the main model from current runtime metadata.** If the identity is unavailable, report it as unverified instead of guessing from a skill name or prior session.
11
11
  2. **Split the work into subtasks and classify each into a lane** (table below).
12
- 3. **Dispatch every task by default — even when the lane's model is you:**
13
- implementation, search, investigation, file reads, verification, tests,
14
- builds, deploys. The commander self-executes only: planning and
15
- decomposition, writing task briefs, reading worker reports and job files
16
- (`out.txt`, `events.jsonl`, inbox records), acceptance judgment, replies to
17
- the operator, git commit/push, and edits to governance files. Read-only
18
- work goes out in advise mode: `triage` for high-volume scans, `long-context`
19
- for large documents, `live-search` for web or X, `hard-judgment` for second
20
- opinions. Editing work uses `--mode work --workdir <repo> --timeout 3600`
21
- or more. Re-verify a worker's claim by reading its attached evidence or by
22
- dispatching a second worker (change `--vendor`); the commander runs no
23
- commands itself. Invalid reasons to skip dispatch: "this lane is mine",
24
- "I am not dispatching so the rule does not apply", "it is only a file
25
- read", "dispatch is slower", "it is one line". Dispatch:
26
- `<repo>/scripts/dispatch.sh [--vendor V] [--mode work] [--workdir DIR] <lane> "<task>"`
12
+ 3. **Delegate every task by default, even to the commander's exact model.**
13
+ Native agents count as delegation; a model match is not permission for
14
+ commander self-execution. Resolve vendor/model/effort separately from
15
+ executor selection. Terminal `auto` without capabilities stays legacy CLI.
16
+ The commander owns planning, task briefs, handoff/completion orchestration,
17
+ reading public results, acceptance, operator replies, git commit/push and
18
+ governance edits. Workers execute the assigned task and never delegate again.
19
+ Read-only work uses advise; edits require `--mode work --workdir <repo>`.
20
+ `<repo>/scripts/dispatch.sh [--executor auto|native|cli] [--native-context FILE] [--vendor V] [--mode work] [--workdir DIR] <lane> "<task>"`
21
+
27
22
  Add `--background` for long tasks; poll with `scripts/jobs.sh status|result <id>`.
28
23
  Use `--thread NAME` when later claude, codex, grok or gemini dispatches
29
24
  must retain earlier context. Threads in 0.33.0 pin vendor, model, effort and
@@ -58,9 +53,59 @@ Doctor remains offline unless the operator explicitly adds `--probe V`; that
58
53
  bounded probe returns metadata only. Use `bin/omnilane benchmark` for a fixed
59
54
  no-call route plan, and add `--run` only when actual advise-mode comparison calls
60
55
  were explicitly requested. Neither command changes routing.
61
- Lanes are fallback chains — dispatch uses the first vendor CLI actually installed,
56
+ Without native context, fallback chains use the first vendor CLI installed,
62
57
  so the same table works with any subset of subscriptions.
63
58
 
59
+ ## Caller-owned native delegation
60
+
61
+ Use explicit current-harness capabilities from the real agent-tool contract:
62
+ active harness/vendor, exact supported model/effort combinations, optional known
63
+ current model, task modes/workdirs, tools, isolation, and lifecycle. Codex
64
+ `collaboration.spawn_agent` has no sandbox/tool/workdir restriction parameters;
65
+ it inherits the parent's tools and filesystem. Advertise `shared-inherited` in
66
+ both request and matching capability row, with empty tool arrays. Treat
67
+ `advise`/`work` and workdir as task intent, not an OS sandbox. Hard isolation is
68
+ CLI-only. Same vendor is not same model. Unknown capabilities do not match.
69
+ Never inspect credentials or infer support from installed CLIs. Explicit
70
+ vendor/model/effort survive native fallback; no next-vendor substitution.
71
+
72
+ ```sh
73
+ omnilane route --executor native --native-context /absolute/capability.json --workdir /absolute/repo hardest-coding "Review the change"
74
+ omnilane jobs --json complete-native JOB_ID /absolute/completion.json
75
+ omnilane jobs --json status JOB_ID
76
+ omnilane jobs --json result JOB_ID
77
+ ```
78
+
79
+ For explicit reuse, the capability must prove the exact existing agent and its
80
+ idle state, and explicitly preserve its existing context. Recheck idle immediately
81
+ before `collaboration.followup_task`; completion must match the reuse strategy,
82
+ agent ID and backend. Do not substitute an unknown inherited model or reuse a busy
83
+ agent. New-agent capacity exhaustion is not success; see `docs/native-executor.md`.
84
+
85
+ Route returns **pending handoff JSON**, not a successful agent run. The host calls
86
+ its own agent tool with the resolved exact model and effort, passes workdir,
87
+ mode, task, and deadline as intent, then ingests the actual agent ID, runtime
88
+ vendor/model/effort/harness/backend, outcome, public result, and evidence. An
89
+ explicit model override uses `fork_turns: "none"` or bounded positive history;
90
+ never combine a model override with `fork_turns: "all"`. An unknown caller
91
+ current model may be omitted when the route explicitly selects an exact model
92
+ declared by the matching capability row. Never report completion before
93
+ ingestion. Native is not a shell executable. The host
94
+ passes **no nested delegation** to native workers; shell workers retain their
95
+ depth guard. Workers do not create handoffs or call agent-spawn tools.
96
+
97
+ Forced CLI retains external workers. Auto explains its CLI fallback reason;
98
+ forced native rejects missing/incompatible capability. CLI sessions (background,
99
+ live, named threads, explicit single-shot), durable/multi-round work,
100
+ vote/arbitration, sysops and unsupported isolation remain CLI-only. The native
101
+ deadline is host-enforced, not a shell watchdog. Native cancellation changes
102
+ pending state without PID signals; the host separately stops any spawned agent.
103
+ Native goal-loop, retry, mailbox, CLI wait and managed-block sync are not included.
104
+
105
+ See [native protocol](../../docs/native-executor.md) for strict schemas, terminal
106
+ examples, lifecycle and public-data boundaries. The parent alone backs up and
107
+ syncs the host's managed `~/.codex/AGENTS.md` block after review.
108
+
64
109
  ## Lanes (defaults; see routing.yaml for the live values)
65
110
 
66
111
  Each lane's **backup** is the next candidate in its `routing.yaml` chain —
@@ -68,26 +113,31 @@ what dispatch picks when the first-choice vendor CLI is not installed.
68
113
 
69
114
  | Lane | First choice | Backup | When |
70
115
  |---|---|---|---|
71
- | hardest-coding | Claude Fable 5.1 (xhigh) | GPT-5.6 Sol (xhigh) → Grok 4.6 → Gemini 3.7 Flash (High) | Hardest implementation, deep root-cause debug, correctness-critical edits |
72
- | bulk-mechanical | GPT-5.6 Sol (high) | Gemini 3.7 Flash (High) → Claude Sonnet 5 (high) | Refactors, migrations, tests, review sweeps — mechanical endurance |
73
- | triage | GPT-5.6 Luna (high) | Gemini 3.7 Flash (Low) → Claude Haiku 4.5 | High-volume scans, first-pass filtering |
74
- | hard-judgment | Claude Opus 5 (xhigh) | GPT-5.6 Sol (max) → Grok 4.6 | Architecture arbitration, deep reasoning, second opinions |
75
- | taste-final | Claude Fable 5.1 (high) | GPT-5.6 Sol (max) → Grok 4.6 → Gemini 3.7 Flash (High) | User-facing prose, prompt/doc polish, Chinese phrasing, style arbitration |
76
- | consult | GPT-5.6 Sol (max) | Claude Fable 5.1 (high) → Grok 4.6 → Gemini 3.7 Flash (High) | Direct named-model consultation; always keep `--vendor` |
77
- | ui-draft | GPT-5.6 Sol (xhigh) | Claude Fable 5.1 (high) → Gemini 3.7 Flash (High) | UI drafts only WITH a design system / reference images; open-ended visual taste goes to taste-final |
78
- | long-context | Gemini 3.7 Flash (Medium) | GPT-5.6 Terra (max) → Claude Opus 5 (medium) | Long-context synthesis ordered on AA-LCR, then cost and throughput |
79
- | fast-agentic | Gemini 3.7 Flash (Medium) | GPT-5.6 Luna (high) → Claude Haiku 4.5 | Fast multi-step agentic loops, multimodal checks |
80
- | live-search | Grok 4.6 | Gemini 3.7 Flash (High) → Claude Sonnet 5 (high) → off | Realtime X/web search and social context; Flash/Sonnet fall back to their own web-search tools |
81
- | coding-overflow | Grok 4.6 | Gemini 3.7 Flash (High) → Kimi K3 → Qwen3 Coder Plus → OpenCode | Codex-quota relief valve for mid-tier coding; verify factual claims |
116
+ | hardest-coding | Claude Fable 5.1 (max) | GPT-6 Astra (xhigh) → Grok 4.6 → Gemini 3.8 Flash (High) | Hardest implementation, deep root-cause debug, correctness-critical edits |
117
+ | bulk-mechanical | GPT-5.6 Sol (high) | Gemini 3.8 Flash (High) → Claude Sonnet 5 (high) | Refactors, migrations, tests, review sweeps — mechanical endurance |
118
+ | triage | GPT-5.6 Luna (high) | Gemini 3.8 Flash (Low) → Claude Haiku 4.5 | High-volume scans, first-pass filtering |
119
+ | hard-judgment | Claude Fable 5.1 (xhigh) | GPT-6 Astra (xhigh) → Grok 4.6 | Architecture arbitration, deep reasoning, second opinions |
120
+ | taste-final | Claude Fable 5.1 (xhigh) | GPT-6 Astra (xhigh) → Grok 4.6 → Gemini 3.8 Flash (High) | User-facing prose, prompt/doc polish, Chinese phrasing, style arbitration |
121
+ | consult | GPT-6 Astra (xhigh) | Claude Fable 5.1 (xhigh) → Grok 4.6 → Gemini 3.8 Flash (Medium) | Direct named-model consultation; always keep `--vendor` |
122
+ | ui-draft | GPT-5.6 Sol (high) | Claude Fable 5.1 (xhigh) → Gemini 3.8 Flash (High) | UI drafts only WITH design system / reference images; open-ended taste goes taste-final |
123
+ | long-context | Gemini 3.8 Flash (Medium) | GPT-5.6 Terra (max) → Claude Opus 5 (medium) | Long-context synthesis; context size alone is not a quality result |
124
+ | fast-agentic | Gemini 3.8 Flash (Low) | GPT-5.6 Luna (high) → Claude Haiku 4.5 | Fast multi-step agentic loops and multimodal checks |
125
+ | live-search | Grok 4.6 | Gemini 3.8 Flash (High) → Claude Sonnet 5 (high) → off | Realtime X/web search; non-Grok fallbacks provide generic web search, not equivalent X context |
126
+ | coding-overflow | Grok 4.6 | Gemini 3.8 Flash (High) → Kimi K3 → Qwen3 Coder Plus → OpenCode → off | Explicit Codex-quota relief; no automatic cross-vendor retry after provider failure |
82
127
  | arbitrate | off (opt-in vote panel) | — | Disabled by default. Enable with `arbitrate: vote codex,claude,grok -` in routing.local.yaml or via the configurator (any 1-4 voters). One quota hit PER VOTER PER ROUND; you chair: read the opinions and own the decision. Effort field 2 = debate round (voters rebut each other) |
83
128
 
84
- Claude Fable 5.1 (`claude-fable-5-1`) is in the taste and hardest-coding
85
- defaults because it leads Opus 5 on every Artificial Analysis axis at the same
86
- effort. It is not in bulk or triage because it prices at twice Opus 5 per token
87
- and consumes the most subscription quota per turn. Opus 5 now leads
88
- hard-judgment on a per-cost, lower-hallucination basis and remains selectable
89
- everywhere via `~/.omnilane/routing.local.yaml`, for example, to bring Fable
90
- back: `hard-judgment: claude claude-fable-5-1 xhigh`.
129
+ Claude Fable 5.1 leads the current hardest-coding, hard-judgment, and
130
+ taste-final defaults at task-specific max/xhigh efforts. GPT-6 Astra is the
131
+ Codex-family fallback and independent-review path. Fable max is the quality-first prompt-level controller. Opus high/xhigh is
132
+ the balanced controller/independent-review option; Astra is the existing-
133
+ Codex-quota backup/reviewer. These are role recommendations, not a lane or
134
+ automatic selector. Opus remains explicitly selectable and in long-context
135
+ fallback.
136
+
137
+ Astra defaults to `xhigh` on the high-difficulty lanes. For an explicitly needed
138
+ upgrade, use `--vendor codex --effort max`; no automatic risk classification or
139
+ failure-triggered effort escalation is added. AA API task costs do not prove
140
+ subscription-quota savings.
91
141
 
92
142
  ## Natural-language consultation
93
143
 
@@ -107,15 +157,19 @@ Users may speak normally; they do not need lane names.
107
157
  | Alias | Vendor | Model | Effort |
108
158
  |---|---|---|---|
109
159
  | Opus | claude | claude-opus-5 | high |
110
- | Fable 5.1 | claude | claude-fable-5-1 | high |
160
+ | Fable 5.1 | claude | claude-fable-5-1 | xhigh |
111
161
  | Sonnet | claude | claude-sonnet-5 | high |
112
162
  | Haiku | claude | claude-haiku-4-5 | - |
113
- | Sol | codex | gpt-5.6-sol | max |
163
+ | Sol | codex | gpt-5.6-sol | high |
114
164
  | Terra | codex | gpt-5.6-terra | max |
115
165
  | Luna | codex | gpt-5.6-luna | high |
166
+ | Astra | codex | gpt-6-astra | xhigh |
116
167
  | Grok 4.6 | grok | grok-4.6 | - |
117
168
  | Gemini 3.1 Pro | gemini | Gemini 3.1 Pro (High) | - |
118
- | Gemini 3.7 Flash | gemini | Gemini 3.7 Flash (High) | - |
169
+ | Gemini 3.8 Flash High | gemini | gemini-3.8-flash-high | - |
170
+ | Gemini 3.8 Flash Medium | gemini | gemini-3.8-flash-medium | - |
171
+ | Gemini 3.8 Flash Low | gemini | gemini-3.8-flash-low | - |
172
+ | Gemini 3.7 Flash | gemini | gemini-3.7-flash-high | - |
119
173
  | Kimi | kimi | kimi-k3 | - |
120
174
  | Qwen | qwen | qwen3-coder-plus | - |
121
175
  | OpenCode | opencode | provider/model form, or `-` for its own default | - |
@@ -152,14 +206,27 @@ dispatch stay in this skill and the CLI. Manage the local board with
152
206
 
153
207
  ## Job lifecycle defaults
154
208
 
155
- - **Completion inbox**: with the Claude Code plugin's hooks installed, a
156
- finished `--background` job is delivered into the foreman's next prompt by
157
- the bundled `UserPromptSubmit` hook, so do not poll for it. Outside Claude
158
- Code, block on `scripts/jobs.sh wait <id> [--timeout N]` instead.
159
- - **Live mailbox**: a `--background` dispatch to Claude or Gemini is a
160
- resident worker. Send follow-up instructions with `scripts/jobs.sh send <id>
161
- "<text>"` and end it with `scripts/jobs.sh close <id>`; other vendors (and
162
- `--single-shot`) run one-shot. Do not use a mailbox for fire-and-forget work.
209
+ - **Active completion (Codex)**: after background CLI dispatch, use
210
+ `scripts/completion-wakeup.py prepare` with the actual controller app thread,
211
+ host, unique run ID and exact job allowlist. Use the returned handoff with the
212
+ app `automation_update` heartbeat tool (reuse an existing monitor), then record
213
+ the actual registration receipt. Do this before ending a turn with unobserved
214
+ jobs. On callback, `poll`, keep unchanged state quiet, acknowledge delivery,
215
+ inspect public results and verify, then acknowledge acceptance with evidence.
216
+ Pause the real automation and record `closed` after all tracked events are
217
+ handled. Never reuse a historical run ID or infer delivery from registration.
218
+ See `docs/completion-wakeup.md` for exact commands and receipt schemas.
219
+ - **Other completion surfaces**: native agent callbacks provide the host result;
220
+ still ingest and verify it. Claude's `UserPromptSubmit` inbox is passive and
221
+ requires another prompt; it does not wake an idle controller. If no supported
222
+ active callback exists, keep the controller active with `scripts/jobs.sh wait`
223
+ and resume acceptance on return rather than asking the user to check again.
224
+ - **Live mailbox**: Claude and Gemini retain automatic resident background
225
+ sessions for supported modes. Codex and Grok remain single-shot by default;
226
+ explicit `--live` opts in. Grok live requires explicit `--mode sysops` because
227
+ ACP has no enforceable restricted-mode boundary. Send follow-up instructions
228
+ with `scripts/jobs.sh send <id> "<text>"` and finish with
229
+ `scripts/jobs.sh close <id>`. `--single-shot` forces one-shot execution.
163
230
  - **Goal orchestration**: when the next step depends on the previous result,
164
231
  wrap the dispatches in `omnilane goal open "<objective>" --workdir DIR`, then
165
232
  `goal dispatch <goal-id> ...`, `goal note`, `goal status`, `goal close --summary`.
@@ -172,50 +239,119 @@ dispatch stay in this skill and the CLI. Manage the local board with
172
239
 
173
240
  - **Dispatch in `advise` mode by default** (read-only worker). Use `--mode work`
174
241
  only when the worker must edit files, and give it an explicit `--workdir`.
175
- - **`--mode sysops`** is `work` minus the vendor sandbox, for service
176
- operations the sandbox denies (launchctl, system daemons). Codex runs with
177
- `-s danger-full-access`; other vendors treat it as `work`. Explicit
178
- per-dispatch opt-in only — never a lane default, and the task text must
179
- name the exact service commands the worker is authorized to run.
180
- Codex `work`/`sysops` still needs a git-repo `--workdir` (non-git
181
- directories trip the whole-job fuse).
242
+ - **Mode contract**: `advise` is read-only with supported native web tools;
243
+ `work` confines file/command changes to explicit `--workdir` and disables
244
+ agent-tool networking, not the model connection. Codex and Claude have
245
+ distinct policies for these modes. Agy 1.1.27 work has bounded new/resume
246
+ acceptance with native sandboxed commands and four validated tools; external
247
+ temp/cache reads are also restricted. A separate two-turn work live/FIFO
248
+ check passed readback, outside-write denial and normal close. macOS Grok work remains blocked because native child-network
249
+ isolation is Linux-only. Do not turn gaps into sysops implicitly or claim
250
+ every provider/mode/session path has passed the runtime matrix.
251
+ - **Agy work tools** are `view_file`, `write_to_file`, `run_command`, and `finish`.
252
+ The native `commandExecutionPolicy: sandbox`, `--sandbox`, and
253
+ `proceed-in-sandbox` policy permits tested workspace edits/builds and denies
254
+ tested outside writes, shell networking and explicit unsandboxed execution.
255
+ Settings are rewritten explicitly before each start/resume: native omission
256
+ of false/empty fields has not been proven default-equivalent. Workspace-local
257
+ caches and the verified empty owned policy directory remain; external cached
258
+ dependencies may be inaccessible, and the tested successful C build still
259
+ emitted an xcrun default-cache denial warning. See the dated capability notes
260
+ for the exact evidence boundary; complete effective SBPL was not captured.
261
+ - **`--mode sysops`** explicitly selects unrestricted native policies for
262
+ Codex, Claude, Grok, and Agy; it is not an alias for work. It is a per-dispatch
263
+ opt-in, never a lane default, and task text must name the allowed operations.
264
+ Codex `work`/`sysops` supports non-Git directories through
265
+ `--skip-git-repo-check`. Without an existing whole-job timeout, dispatch
266
+ adds one when its supervisor is available; otherwise it warns and retains
267
+ the per-call watchdog path.
268
+ The CLI defaults an omitted `--workdir` to the caller’s current directory;
269
+ task briefs must still specify it explicitly. The MCP work interface
270
+ separately requires an explicit `workdir`.
271
+ - **Grok advise web tools** use internal `web_search` / `web_fetch` selectors,
272
+ while permission rules keep their native `WebSearch` / `WebFetch` class names.
273
+ The complete single-shot `plain` path has real search, fetched-page, and native
274
+ denied-write evidence on Grok 1.0.13. MCP readiness is job-local because this
275
+ mode denies MCPTool; hooks and their security checks remain enabled. A caller
276
+ supplied nonempty `CONTEXT_MODE_MCP_SENTINEL_DIR` is a conflict and stops before provider
277
+ startup rather than being overwritten. Do not extend this result to restricted
278
+ live or macOS work.
182
279
  - **Every dispatched task states acceptance criteria and the exact verification
183
280
  command.** Do not accept "done" without evidence.
184
281
  - **No nested dispatch**: workers must not fan out again (enforced via
185
282
  `OMNILANE_DEPTH`). Escalate back to the main loop instead.
186
283
  - **Same-directory codex dispatches are serialized automatically** (lock);
187
284
  do not try to parallelize them yourself.
188
- - Escalate without asking: two failed attempts on a lane → move one lane up
189
- (triage → bulk-mechanical → hardest-coding).
285
+ - After two failed attempts, reassess scope and retry only an eligible exact configuration.
286
+ An upward AA move requires the human to take over; a model cannot approve its own uplift.
190
287
  - Vendor quota exhausted (429 / "stream disconnected" / usage-limit message):
191
288
  send mid-tier coding through coding-overflow instead; never silently downgrade
192
289
  hardest-coding — wait or escalate to the user.
193
290
 
194
291
  ## Per-model notes (apply the row matching YOUR main model)
195
292
 
196
- - **Claude Fable 5.1 main**: taste finalization and the hardest coding are
197
- yours; hard judgment now defaults to Opus 5. Dispatch bulk work to Sol high,
198
- hard judgment to Opus 5 xhigh, and long-context or fast loops to Gemini 3.7
199
- Flash.
200
- - **Claude Opus 5 main**: hard judgment is now yours by default. Taste-final
201
- remains Fable's; use local overrides when Opus's lower hallucination rate or
202
- price is preferred there.
203
- - **Claude Sonnet main**: coordination/tools/mid-tier coding only, plus
204
- fallback duty in bulk-mechanical and live-search; never self-assign top
205
- judgment or hardest implementation.
206
- - **GPT Sol main**: hardest coding + hard judgment are yours (use max for
207
- judgment turns, xhigh for coding); cross to taste-final for style calls.
208
- - **GPT Terra main**: long-context Codex fallback work is yours at max;
209
- bulk-mechanical now defaults to Sol high, and genuinely hardest pieces
210
- escalate to Sol xhigh.
211
- - **Grok 4.6 main**: live-search and coding overflow are yours, plus fallback
212
- duty in hardest-coding, hard-judgment, and taste-final; its measured
213
- hallucination rate is the lowest among the frontier rows, but still verify
214
- every API signature and cited fact before shipping.
215
- - **Gemini 3.7 Flash main**: long-context and fast agentic/multimodal loops
216
- are yours at the lane's configured effort; bulk and overflow use the high
217
- row, plus fallback duty (High) in hardest-coding, taste-final, ui-draft, and
218
- live-search. Never self-assign top judgment.
219
- - **Gemini 3.1 Pro main**: it remains directly selectable, but the default
220
- long-context lane now prefers Gemini 3.7 Flash on LCR, cost, and throughput;
221
- route hardest coding and judgment to the stronger Codex and Claude lanes.
293
+ These notes never expand the commander's reserved self-execution scope. If the verified main model has no matching row, use the configured lane table under the current user request and `rules.d/60`; do not assume the nearest older model is equivalent or silently override vendor/model/effort. If a required capability or explicit model choice is unresolved, report that exact gap before dispatch rather than inventing a fallback.
294
+
295
+ - **Claude Fable 5.1 main**: recommended prompt-level controller for
296
+ quality-sensitive work (not a lane or automatic selector). Hardest coding
297
+ uses max; judgment and taste use xhigh. Dispatch bulk work to Sol high,
298
+ long/fast work to Gemini 3.8 Flash, and use Astra as an independent Codex
299
+ review path.
300
+ - **Claude Opus 5 main**: balanced prompt-level controller and independent
301
+ reviewer when explicitly selected (`high`, or `xhigh` for deeper review),
302
+ plus Claude long-context fallback. This is a role/opt-in choice, not a new
303
+ lane or the current hard-judgment default.
304
+ - **Claude Sonnet main**: coordination/tools/mid-tier coding only, plus fallback
305
+ duty in bulk-mechanical and live-search; never self-assign top judgment or
306
+ hardest implementation.
307
+ - **GPT Astra main**: controller backup and independent reviewer. Resolve the
308
+ actual caller effort first; select only an exact target at or below its ceiling.
309
+ An explicit higher-effort request does not bypass model-level AA policy.
310
+ - **GPT Sol main**: mechanical work and constrained UI drafts only within the
311
+ exact effective ceiling; return higher-score needs to the operator.
312
+ - **GPT Terra main**: long-context and mechanical work only within the exact
313
+ effective ceiling; do not infer eligibility from the Terra family label.
314
+ - **GPT Luna main**: high-volume triage is delegated at high; do not promote its
315
+ low price into correctness-critical or controller work.
316
+ - **Grok 4.6 main**: live-search and coding-overflow are yours, plus fallback
317
+ duty in hard lanes. Grok effort remains ignored; verify API signatures and
318
+ cited facts before shipping.
319
+ - **Gemini 3.8 Flash main**: long-context uses medium, fast-agentic and triage
320
+ use low, and bulk/overflow/web fallbacks use high. Do not infer visual taste
321
+ or controller authority from agent/coding benchmarks.
322
+ - **Gemini 3.1 Pro main**: remains directly selectable, but is not promoted by
323
+ this refresh; route hard coding and judgment to the stronger configured lanes.
324
+
325
+ ## Frozen exact-AA downward gate
326
+
327
+ All lane and per-model preferences above are subordinate to this gate, including
328
+ explicit vendor/model/effort requests. Supply `--caller-context /absolute/context.json`
329
+ with schema_version=1, snapshot_id, kind=model, caller containing exact vendor/model/
330
+ effort/reasoning/fallback, and inherited_ceiling. Effective ceiling is the minimum of
331
+ that exact frozen score and the inherited ceiling. Targets at or below it are allowed;
332
+ unknown identities and unresolved request-selector mappings fail closed. No family,
333
+ displayed grade, highest-effort assumption, retry or fallback grants an uplift.
334
+
335
+ A `--transport-overlay /absolute/overlay.json` may prove a small set of host-local
336
+ request selectors using exact identities and hashed local contract evidence. It does
337
+ not change frozen AA scores or certify upstream provider identity. The explicit
338
+ `--operator-asserted-human` exemption is cooperative operator metadata, not automatic
339
+ model detection or OS authentication; model callers must not assert it for themselves.
340
+
341
+ CLI jobs atomically save an original authorizer, exact child caller context, decision,
342
+ and registry snapshot. Provider processes receive the child identity, not the parent's.
343
+ Retries preserve original target config, check stored hashes, and intersect the
344
+ current exact caller score/inherited ceiling with the original authorizer ceiling.
345
+ Use `omnilane jobs retry ID --caller-context FILE`; missing current identity fails
346
+ closed, and a model retry never inherits an earlier human exemption. Native handoffs carry the same decision plus a job-owned
347
+ `worker_contract.caller_context_path`; pass that context to the native child together
348
+ with the no-nested-dispatch requirement. `OMNILANE_DEPTH` remains an independent guard.
349
+
350
+ A background job or PENDING native handoff is not completion. Observe its terminal
351
+ result and acceptance evidence before closing the controller task. Completion wakeup
352
+ availability must be separately verified; never claim delivery from scheduling alone.
353
+
354
+ For Gemini `model_id_encoded_effort` selectors, the proven native model ID encodes
355
+ the AA effort. An absent parity effort or the matching effort is accepted; a conflicting
356
+ parity effort is rejected. Do not represent a discarded EFFORT parameter as an active
357
+ provider setting. Frozen reasoning=`unspecified` remains a literal, not a wildcard.