omnilane 0.31.0 → 0.32.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -8,7 +8,7 @@
8
8
  # consult is a multi-vendor direct-target chain. configure.sh intentionally
9
9
  # skips it because that menu writes one candidate per lane. If overriding it,
10
10
  # retain every vendor you want to address by name:
11
- # consult: codex gpt-5.6-sol max | claude claude-opus-5 high | grok grok-4.5 - | gemini "Gemini 3.1 Pro (High)" -
11
+ # consult: codex gpt-5.6-sol max | claude claude-fable-5-1 high | grok grok-4.6 - | gemini "Gemini 3.7 Flash (High)" -
12
12
 
13
13
  # ── Starter profiles ─────────────────────────────────────────────
14
14
  # Uncomment ONE block that matches what you actually subscribe to.
@@ -26,10 +26,10 @@
26
26
  # live-search: off - -
27
27
  # coding-overflow: off - -
28
28
 
29
- # Profile: Claude Code (Fable 5) as the main loop — let Fable keep judgment/taste,
29
+ # Profile: Claude Code (Fable 5.1) as the main loop — let Fable keep judgment/taste,
30
30
  # push only coding volume out to Codex.
31
- # hard-judgment: claude claude-fable-5 high
32
- # taste-final: claude claude-fable-5 high
31
+ # hard-judgment: claude claude-fable-5-1 high
32
+ # taste-final: claude claude-fable-5-1 high
33
33
 
34
34
  # Profile: Codex-heavy (Sol main) — keep the hard lanes on Codex, Claude for taste.
35
35
  # taste-final: claude claude-opus-5 high
package/routing.yaml CHANGED
@@ -8,39 +8,37 @@
8
8
  # same table degrades gracefully when you only subscribe to one or two vendors.
9
9
  # Override any line in ~/.omnilane/routing.local.yaml (same format; local wins).
10
10
  # Every benchmark score, price and throughput figure behind these orderings lives in
11
- # docs/model-capabilities-2026-07.md, with the date it was retrieved. The comments below
11
+ # docs/model-capabilities-2026-09.md, with the date it was retrieved. The comments below
12
12
  # deliberately carry no numbers: they state WHY a lane is ordered the way it is, which
13
13
  # stays true for months, while the numbers move every few weeks. Change an ordering and
14
14
  # you update the doc; a figure going stale should never need a routing-table edit.
15
- # (Audited 2026-07-12; re-audited 2026-07-25, 2026-08-02 and 2026-08-03.)
16
- # defaults follow Artificial Analysis data, 2026-07
15
+ # (Audited 2026-07-12; re-audited 2026-07-25, 2026-08-02, 2026-08-03, and 2026-09-02.)
16
+ # defaults follow Artificial Analysis data, 2026-09
17
17
  # snapshot. Verified against AA site records + vendor pricing pages: Intelligence &
18
18
  # Coding indexes and 7:2:1 blended prices all match (AA field price1mBlended7To2To1);
19
19
  # coding cost-per-task is chart-only (not independently reconstructed). Prices are
20
20
  # standard short-context API tier — on subscription CLIs treat $ as relative ranking.
21
21
  # Your own job outcomes (~/.omnilane/jobs/) outrank these priors; edit lanes to match.
22
22
 
23
- hardest-coding: codex gpt-5.6-sol xhigh | claude claude-opus-5 xhigh # ordered on coding capability specifically, not general intelligence. Sol dropped from max to xhigh on 2026-08-03: on AA's per-effort Coding Index, Sol at xhigh outscores both Sol at max and every Claude tier, at a third less cost — max buys overthinking here, not accuracy. xhigh is also Anthropic's documented starting point for coding/agentic work. Keep Sol first for the established Codex harness lane.
24
- bulk-mechanical: codex gpt-5.6-terra max | claude claude-sonnet-5 high | gemini "Gemini 3.6 Flash (High)" - # ordered on endurance per dollar: Terra leads Sonnet 5 on both intelligence and cost per task
25
- triage: codex gpt-5.6-luna medium | gemini "Gemini 3.6 Flash (Low)" - | claude claude-haiku-4-5 - # high-volume scans, ordered on cost per task: Luna is the cheapest model at its intelligence tier by a wide margin
26
- hard-judgment: claude claude-opus-5 xhigh | codex gpt-5.6-sol max # ordered on agentic knowledge work, where Opus 5 leads Sol on AA's benchmarks. xhigh per Anthropic guidance (high is the documented floor for intelligence-sensitive work; max is for correctness-over-cost only) raise to max locally via `omnilane configure set` if your workload needs it.
27
- taste-final: claude claude-opus-5 high | codex gpt-5.6-sol max # user-facing prose, prompt/doc polish, Chinese phrasing, style arbitration
28
- consult: codex gpt-5.6-sol max | claude claude-opus-5 high | grok grok-4.5 - | gemini "Gemini 3.1 Pro (High)" - # direct named-model consultation; use --vendor to prevent fallback
29
- ui-draft: codex gpt-5.6-sol xhigh | claude claude-opus-5 high # only with a design system / reference images; open-ended visual taste -> taste-final
30
- long-context: gemini "Gemini 3.1 Pro (High)" - | codex gpt-5.6-sol high | claude claude-opus-5 high # all have 1M context; ordered on AA-LCR, which scores exactly this lane's work — extracting and synthesising across long documents — and where Gemini leads both fallbacks. Corrected 2026-08-03: this comment used to send multi-hop synthesis to the Claude candidate on second-hand prior-generation figures, and current first-party per-effort data reverses that, so the two fallbacks swapped. Caveat in docs: AA-LCR runs at 10k-100k tokens, so nothing here settles behaviour at a full 1M
31
- fast-agentic: codex gpt-5.6-luna max | gemini "Gemini 3.6 Flash (High)" - # fast multi-step tool loops. Reordered 2026-08-03: Luna leads Flash on agentic benchmarks AND costs a fraction as much per task, so Flash's remaining edge is raw throughput alone. Keep Flash first only if your loops are latency-bound. Both take image input, so the lane's multimodal checks are unaffected
32
- live-search: grok grok-4.5 - | off # native X/web search lane; no real substitute
33
- coding-overflow: grok grok-4.5 - | kimi kimi-k3 - | qwen qwen3-coder-plus - | opencode - - | off # codex-quota relief valve: mid-tier coding; Grok 4.5 is a capable mid-tier coder but AA measures a high hallucination rate verify every factual claim it ships. qwen3-coder-plus = 2025-09-23 snapshot alias (Qwen 3.6 Plus exists; re-evaluate before swapping). kimi/qwen model fields are CLI aliases adjust to your login. opencode "-" model = its own configured default.
34
- arbitrate: off - - # opinion panel is OPT-IN: it costs one call per voter per round.
23
+ hardest-coding: claude claude-fable-5-1 xhigh | codex gpt-5.6-sol xhigh # ordered on coding capability: the leading Claude tier wins both coding components; Sol remains the established Codex-harness value fallback
24
+ bulk-mechanical: codex gpt-5.6-sol high | gemini "Gemini 3.7 Flash (High)" - | claude claude-sonnet-5 high # ordered on endurance per dollar: Sol dominates Terra within Codex; current Flash is the faster, cheaper middle fallback; Sonnet preserves subscription quota
25
+ triage: codex gpt-5.6-luna high | gemini "Gemini 3.7 Flash (Low)" - | claude claude-haiku-4-5 - # ordered on cost per task at usable intelligence: Luna high buys a meaningful quality lift cheaply; Flash and Haiku are low-cost cross-vendor fallbacks
26
+ hard-judgment: claude claude-fable-5-1 xhigh | codex gpt-5.6-sol max | grok grok-4.6 - # ordered on agentic knowledge work: Fable leads the Claude field; Sol stays ahead of Grok because Grok effort is ignored and its reproduced row is unknown
27
+ taste-final: claude claude-fable-5-1 high | codex gpt-5.6-sol max # ordered on prose and polish: Fable leads Opus on intelligence and factual breadth; Sol is the cross-vendor fallback
28
+ consult: codex gpt-5.6-sol max | claude claude-fable-5-1 high | grok grok-4.6 - | gemini "Gemini 3.7 Flash (High)" - # direct named-model chain uses the strongest current Claude and Flash slots; keep --vendor to prevent fallback
29
+ ui-draft: codex gpt-5.6-sol xhigh | claude claude-fable-5-1 high # ordered for drafts with a design system or reference images: Sol leads measured multimodal and coding evidence; Fable follows for polish
30
+ long-context: gemini "Gemini 3.7 Flash (Medium)" - | codex gpt-5.6-terra max | claude claude-opus-5 medium # ordered on long-context reasoning, then cost and throughput: Flash leads; Terra matches its long-context result; Opus is the cheaper Claude fallback
31
+ fast-agentic: gemini "Gemini 3.7 Flash (Medium)" - | codex gpt-5.6-luna high # ordered on interactive tool-loop latency: Flash gives up little agentic quality for far faster first output; Luna high is the low-latency Codex fallback
32
+ live-search: grok grok-4.6 - | off # native X and web search lane; no real substitute
33
+ coding-overflow: grok grok-4.6 - | gemini "Gemini 3.7 Flash (High)" - | kimi kimi-k3 - | qwen qwen3-coder-plus - | opencode - - | off # coding relief ordered by capability and value: Grok has the lowest frontier hallucination rate; Flash is the cheapest strong coder here; revisit the best-value Qwen tier when its CLI alias can be verified
34
+ arbitrate: off - - # opinion panel remains opt-in because each voter and round consumes quota
35
35
  # Enable: `arbitrate: vote codex,claude,grok -` (any 1-4 of codex/claude/grok/gemini)
36
36
  # Debate round (each voter rebuts the others): set the effort field to 2.
37
37
  # Custom gate: `arbitrate: exec /path/to/script -`
38
- # Claude Fable 5 (claude-fable-5) is deliberately absent from the defaults: the top Claude tier
39
- # is usually the MAIN LOOP itself, not a dispatched worker, and it prices at twice Opus 5.
40
- # This is a cost / guardrail / main-loop policy choice, NOT a capability verdict AA and
41
- # Epoch AI disagree on which of the two leads general intelligence and call it effectively a
42
- # tie, while Opus 5 leads clearly on agentic knowledge work at a lower cost per task. Fable 5
43
- # does keep the lead on factual breadth, so name it explicitly for recall-heavy consults. If you want
44
- # to route to it anyway, pick it in the configurator or set e.g.
45
- # taste-final: claude claude-fable-5 high
46
- # in ~/.omnilane/routing.local.yaml.
38
+ # Claude Fable 5.1 is in the judgment, taste, and hardest-coding defaults because
39
+ # it leads Opus 5 on every Artificial Analysis axis at the same effort.
40
+ # It is not in bulk or triage: it prices at twice Opus 5 per token and consumes
41
+ # the most subscription quota per turn. Opus 5 remains the lower-hallucination,
42
+ # lower-price Claude choice and can return to any lane via
43
+ # ~/.omnilane/routing.local.yaml, for example:
44
+ # hard-judgment: claude claude-opus-5 xhigh
@@ -151,10 +151,10 @@ esac
151
151
  # Dynamic/API catalogs stay curated — "c" always accepts an exact model ID.
152
152
  CODEX_MODELS=("gpt-5.6" "gpt-5.6-sol" "gpt-5.6-terra" "gpt-5.6-luna" "gpt-5.5" "gpt-5.4" "gpt-5.4-mini" "gpt-5.3-codex-spark")
153
153
  CODEX_EFFORTS=("xhigh" "max" "ultra" "high" "medium" "low" "minimal" "none")
154
- CLAUDE_MODELS=("default" "best" "fable" "opus" "sonnet" "haiku" "opus[1m]" "sonnet[1m]" "opusplan" "claude-fable-5" "claude-opus-5" "claude-sonnet-5" "claude-opus-4-8" "claude-opus-4-7" "claude-opus-4-6" "claude-opus-4-5-20251101" "claude-sonnet-4-6" "claude-sonnet-4-5-20250929" "claude-haiku-4-5" "claude-haiku-4-5-20251001")
154
+ CLAUDE_MODELS=("default" "best" "fable" "opus" "sonnet" "haiku" "opus[1m]" "sonnet[1m]" "opusplan" "claude-fable-5" "claude-fable-5-1" "claude-opus-5" "claude-sonnet-5" "claude-opus-4-8" "claude-opus-4-7" "claude-opus-4-6" "claude-opus-4-5-20251101" "claude-sonnet-4-6" "claude-sonnet-4-5-20250929" "claude-haiku-4-5" "claude-haiku-4-5-20251001")
155
155
  CLAUDE_EFFORTS=("max" "xhigh" "high" "medium" "low" "-")
156
- GEMINI_MODELS=("gemini-3.6-flash-high" "gemini-3.6-flash-medium" "gemini-3.6-flash-low" "gemini-3.5-flash-high" "gemini-3.5-flash-medium" "gemini-3.5-flash-low" "gemini-3.1-pro-high" "gemini-3.1-pro-low" "claude-sonnet-4-6" "claude-opus-4-6-thinking" "gpt-oss-120b-medium")
157
- GROK_MODELS=("grok-4.5" "headroom-grok-build" "grok-4.3-official")
156
+ GEMINI_MODELS=("gemini-3.7-flash-high" "gemini-3.7-flash-medium" "gemini-3.7-flash-low" "gemini-3.6-flash-high" "gemini-3.6-flash-medium" "gemini-3.6-flash-low" "gemini-3.1-pro-high" "gemini-3.1-pro-low" "claude-sonnet-4-6" "claude-opus-4-6-thinking" "gpt-oss-120b-medium")
157
+ GROK_MODELS=("grok-4.6" "headroom-grok-build" "grok-4.3-official")
158
158
  KIMI_MODELS=("kimi-k3" "kimi-k2.7-code" "kimi-k2.5")
159
159
  QWEN_MODELS=("qwen3.7-max" "qwen3.7-plus" "qwen3.6-plus" "qwen3.5-plus" "qwen3-max-2026-01-23" "qwen3-coder-next" "qwen3-coder-plus" "qwen3-coder-flash")
160
160
  # OpenCode models use provider/model form; OpenRouter models use catalog slugs.
@@ -31,9 +31,9 @@ trap cleanup_temp_files EXIT
31
31
  voter_spec() { # vendor -> "model<TAB>effort"
32
32
  case "$1" in
33
33
  codex) printf 'gpt-5.6-sol\thigh' ;;
34
- claude) printf 'claude-opus-5\thigh' ;;
35
- gemini) printf 'Gemini 3.1 Pro (High)\t-' ;;
36
- grok) printf 'grok-4.5\t-' ;;
34
+ claude) printf 'claude-fable-5-1\thigh' ;;
35
+ gemini) printf 'Gemini 3.7 Flash (High)\t-' ;;
36
+ grok) printf 'grok-4.6\t-' ;;
37
37
  *) return 1 ;;
38
38
  esac
39
39
  }
@@ -50,28 +50,26 @@ what dispatch picks when the first-choice vendor CLI is not installed.
50
50
 
51
51
  | Lane | First choice | Backup | When |
52
52
  |---|---|---|---|
53
- | hardest-coding | GPT-5.6 Sol (xhigh) | Claude Opus 5 (xhigh) | Hardest implementation, deep root-cause debug, correctness-critical edits |
54
- | bulk-mechanical | GPT-5.6 Terra (max) | Claude Sonnet 5 (high) | Refactors, migrations, tests, review sweeps — mechanical endurance |
55
- | triage | GPT-5.6 Luna (medium) | Gemini 3.6 Flash (Low) | High-volume scans, first-pass filtering |
56
- | hard-judgment | Claude Opus 5 (xhigh) | GPT-5.6 Sol (max) | Architecture arbitration, deep reasoning, second opinions |
57
- | taste-final | Claude Opus 5 (high) | GPT-5.6 Sol (max) | User-facing prose, prompt/doc polish, Chinese phrasing, style arbitration |
58
- | consult | Explicit named vendor/model | (no fallback) | Direct natural-language consultation; always keep `--vendor` |
59
- | ui-draft | GPT-5.6 Sol (xhigh) | Claude Opus 5 (high) | UI drafts only WITH a design system / reference images; open-ended visual taste goes to taste-final |
60
- | long-context | Gemini 3.1 Pro (High) | GPT-5.6 Sol (high) | 1M-token synthesis; Pro is agentic-capable, while fast repeated loops prefer Flash on speed/cost |
61
- | fast-agentic | GPT-5.6 Luna (max) | Gemini 3.6 Flash (High) | Fast multi-step agentic loops, multimodal checks |
62
- | live-search | Grok 4.5 | — (off) | Realtime X/web search and social context |
63
- | coding-overflow | Grok 4.5 | Kimi K3 → Qwen3 Coder Plus → OpenCode | Codex-quota relief valve for mid-tier coding; verify factual claims |
53
+ | hardest-coding | Claude Fable 5.1 (xhigh) | GPT-5.6 Sol (xhigh) | Hardest implementation, deep root-cause debug, correctness-critical edits |
54
+ | bulk-mechanical | GPT-5.6 Sol (high) | Gemini 3.7 Flash (High) → Claude Sonnet 5 (high) | Refactors, migrations, tests, review sweeps — mechanical endurance |
55
+ | triage | GPT-5.6 Luna (high) | Gemini 3.7 Flash (Low) → Claude Haiku 4.5 | High-volume scans, first-pass filtering |
56
+ | hard-judgment | Claude Fable 5.1 (xhigh) | GPT-5.6 Sol (max) → Grok 4.6 | Architecture arbitration, deep reasoning, second opinions |
57
+ | taste-final | Claude Fable 5.1 (high) | GPT-5.6 Sol (max) | User-facing prose, prompt/doc polish, Chinese phrasing, style arbitration |
58
+ | consult | GPT-5.6 Sol (max) | Claude Fable 5.1 (high) → Grok 4.6 → Gemini 3.7 Flash (High) | Direct named-model consultation; always keep `--vendor` |
59
+ | ui-draft | GPT-5.6 Sol (xhigh) | Claude Fable 5.1 (high) | UI drafts only WITH a design system / reference images; open-ended visual taste goes to taste-final |
60
+ | long-context | Gemini 3.7 Flash (Medium) | GPT-5.6 Terra (max) → Claude Opus 5 (medium) | Long-context synthesis ordered on AA-LCR, then cost and throughput |
61
+ | fast-agentic | Gemini 3.7 Flash (Medium) | GPT-5.6 Luna (high) | Fast multi-step agentic loops, multimodal checks |
62
+ | live-search | Grok 4.6 | — (off) | Realtime X/web search and social context |
63
+ | coding-overflow | Grok 4.6 | Gemini 3.7 Flash (High) → Kimi K3 → Qwen3 Coder Plus → OpenCode | Codex-quota relief valve for mid-tier coding; verify factual claims |
64
64
  | arbitrate | off (opt-in vote panel) | — | Disabled by default. Enable with `arbitrate: vote codex,claude,grok -` in routing.local.yaml or via the configurator (any 1-4 voters). One quota hit PER VOTER PER ROUND; you chair: read the opinions and own the decision. Effort field 2 = debate round (voters rebut each other) |
65
65
 
66
- Claude Fable 5 (`claude-fable-5`) is absent from the defaults on purpose: the
67
- top Claude tier is usually the main loop itself, not a dispatched worker, and
68
- it prices at twice Opus 5. This is a cost / guardrail / main-loop policy choice,
69
- not a capability verdict Artificial Analysis calls Opus 5 (61) and Fable 5 (60)
70
- "effectively tied" on the Intelligence Index, but Opus 5 leads AA-Briefcase by
71
- 146 Elo at 20% lower cost per task. Fable 5 keeps the lead on factual breadth
72
- (AA-Omniscience), so name it explicitly for recall-heavy consults. To route to
73
- it anyway, select it in the configurator or override a lane in
74
- `~/.omnilane/routing.local.yaml` (e.g. `taste-final: claude claude-fable-5 high`).
66
+ Claude Fable 5.1 (`claude-fable-5-1`) is in the judgment, taste, and
67
+ hardest-coding defaults because it leads Opus 5 on every Artificial Analysis
68
+ axis at the same effort. It is not in bulk or triage because it prices at twice
69
+ Opus 5 per token and consumes the most subscription quota per turn. Opus 5
70
+ remains the lower-hallucination, lower-price Claude choice and can return to any
71
+ lane via `~/.omnilane/routing.local.yaml`, for example:
72
+ `hard-judgment: claude claude-opus-5 xhigh`.
75
73
 
76
74
  ## Natural-language consultation
77
75
 
@@ -91,15 +89,15 @@ Users may speak normally; they do not need lane names.
91
89
  | Alias | Vendor | Model | Effort |
92
90
  |---|---|---|---|
93
91
  | Opus | claude | claude-opus-5 | high |
94
- | Fable | claude | claude-fable-5 | high |
92
+ | Fable 5.1 | claude | claude-fable-5-1 | high |
95
93
  | Sonnet | claude | claude-sonnet-5 | high |
96
94
  | Haiku | claude | claude-haiku-4-5 | - |
97
95
  | Sol | codex | gpt-5.6-sol | max |
98
96
  | Terra | codex | gpt-5.6-terra | max |
99
- | Luna | codex | gpt-5.6-luna | medium |
100
- | Grok 4.5 | grok | grok-4.5 | - |
101
- | Gemini Pro | gemini | Gemini 3.1 Pro (High) | - |
102
- | Gemini Flash | gemini | Gemini 3.6 Flash (High) | - |
97
+ | Luna | codex | gpt-5.6-luna | high |
98
+ | Grok 4.6 | grok | grok-4.6 | - |
99
+ | Gemini 3.1 Pro | gemini | Gemini 3.1 Pro (High) | - |
100
+ | Gemini 3.7 Flash | gemini | Gemini 3.7 Flash (High) | - |
103
101
  | Kimi | kimi | kimi-k3 | - |
104
102
  | Qwen | qwen | qwen3-coder-plus | - |
105
103
  | OpenCode | opencode | provider/model form, or `-` for its own default | - |
@@ -159,19 +157,24 @@ dispatch stay in this skill and the CLI. Manage the local board with
159
157
 
160
158
  ## Per-model notes (apply the row matching YOUR main model)
161
159
 
162
- - **Claude (Fable/Opus main)**: top judgment and taste are yours, but
163
- implementation still dispatches by default self-execute only reserved
164
- commander items; push all coding volume out to the lanes.
160
+ - **Claude Fable 5.1 main**: hard judgment, taste finalization, and the hardest
161
+ coding are yours. Dispatch bulk work to Sol high and long-context or fast
162
+ loops to Gemini 3.7 Flash.
163
+ - **Claude Opus 5 main**: judgment and taste remain strong self-execute lanes;
164
+ use local overrides when its lower hallucination rate or price is preferred.
165
165
  - **Claude Sonnet main**: coordination/tools/mid-tier coding only; never
166
166
  self-assign top judgment or hardest implementation.
167
167
  - **GPT Sol main**: hardest coding + hard judgment are yours (use max for
168
168
  judgment turns, xhigh for coding); cross to taste-final for style calls.
169
- - **GPT Terra main**: bulk work is yours at max; escalate the genuinely hardest
170
- pieces to Sol instead of grinding.
171
- - **Grok 4.5 main**: mid-tier coding + live-search are yours; verify every API
172
- signature and cited fact before shipping (measured high hallucination rate).
173
- - **Gemini Flash main**: fast agentic/multimodal loops are yours; never
174
- self-assign top judgment.
175
- - **Gemini 3.1 Pro main**: 1M-context synthesis and context-heavy agentic work
176
- are yours. Prefer Gemini Flash for fast repeated tool loops on speed/cost;
177
- route hardest coding and judgment to the stronger codex lanes.
169
+ - **GPT Terra main**: long-context Codex fallback work is yours at max;
170
+ bulk-mechanical now defaults to Sol high, and genuinely hardest pieces
171
+ escalate to Sol xhigh.
172
+ - **Grok 4.6 main**: live-search and coding overflow are yours; its measured
173
+ hallucination rate is the lowest among the frontier rows, but still verify
174
+ every API signature and cited fact before shipping.
175
+ - **Gemini 3.7 Flash main**: long-context and fast agentic/multimodal loops
176
+ are yours at the lane's configured effort; bulk and overflow use the high row.
177
+ Never self-assign top judgment.
178
+ - **Gemini 3.1 Pro main**: it remains directly selectable, but the default
179
+ long-context lane now prefers Gemini 3.7 Flash on LCR, cost, and throughput;
180
+ route hardest coding and judgment to the stronger Codex and Claude lanes.