omnilane 0.31.0 → 0.32.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +2 -2
- package/.claude-plugin/plugin.json +1 -1
- package/CHANGELOG.md +25 -1
- package/README.ja.md +51 -65
- package/README.ko.md +51 -64
- package/README.md +52 -66
- package/README.zh-CN.md +50 -61
- package/README.zh-TW.md +51 -62
- package/VERSION +1 -1
- package/package.json +1 -1
- package/plugin.json +1 -1
- package/routing.local.yaml.example +4 -4
- package/routing.yaml +22 -24
- package/scripts/configure.sh +3 -3
- package/scripts/runners/run-vote.sh +3 -3
- package/skills/omnilane/SKILL.md +40 -37
|
@@ -8,7 +8,7 @@
|
|
|
8
8
|
# consult is a multi-vendor direct-target chain. configure.sh intentionally
|
|
9
9
|
# skips it because that menu writes one candidate per lane. If overriding it,
|
|
10
10
|
# retain every vendor you want to address by name:
|
|
11
|
-
# consult: codex gpt-5.6-sol max | claude claude-
|
|
11
|
+
# consult: codex gpt-5.6-sol max | claude claude-fable-5-1 high | grok grok-4.6 - | gemini "Gemini 3.7 Flash (High)" -
|
|
12
12
|
|
|
13
13
|
# ── Starter profiles ─────────────────────────────────────────────
|
|
14
14
|
# Uncomment ONE block that matches what you actually subscribe to.
|
|
@@ -26,10 +26,10 @@
|
|
|
26
26
|
# live-search: off - -
|
|
27
27
|
# coding-overflow: off - -
|
|
28
28
|
|
|
29
|
-
# Profile: Claude Code (Fable 5) as the main loop — let Fable keep judgment/taste,
|
|
29
|
+
# Profile: Claude Code (Fable 5.1) as the main loop — let Fable keep judgment/taste,
|
|
30
30
|
# push only coding volume out to Codex.
|
|
31
|
-
# hard-judgment: claude claude-fable-5 high
|
|
32
|
-
# taste-final: claude claude-fable-5 high
|
|
31
|
+
# hard-judgment: claude claude-fable-5-1 high
|
|
32
|
+
# taste-final: claude claude-fable-5-1 high
|
|
33
33
|
|
|
34
34
|
# Profile: Codex-heavy (Sol main) — keep the hard lanes on Codex, Claude for taste.
|
|
35
35
|
# taste-final: claude claude-opus-5 high
|
package/routing.yaml
CHANGED
|
@@ -8,39 +8,37 @@
|
|
|
8
8
|
# same table degrades gracefully when you only subscribe to one or two vendors.
|
|
9
9
|
# Override any line in ~/.omnilane/routing.local.yaml (same format; local wins).
|
|
10
10
|
# Every benchmark score, price and throughput figure behind these orderings lives in
|
|
11
|
-
# docs/model-capabilities-2026-
|
|
11
|
+
# docs/model-capabilities-2026-09.md, with the date it was retrieved. The comments below
|
|
12
12
|
# deliberately carry no numbers: they state WHY a lane is ordered the way it is, which
|
|
13
13
|
# stays true for months, while the numbers move every few weeks. Change an ordering and
|
|
14
14
|
# you update the doc; a figure going stale should never need a routing-table edit.
|
|
15
|
-
# (Audited 2026-07-12; re-audited 2026-07-25, 2026-08-02 and 2026-
|
|
16
|
-
# defaults follow Artificial Analysis data, 2026-
|
|
15
|
+
# (Audited 2026-07-12; re-audited 2026-07-25, 2026-08-02, 2026-08-03, and 2026-09-02.)
|
|
16
|
+
# defaults follow Artificial Analysis data, 2026-09
|
|
17
17
|
# snapshot. Verified against AA site records + vendor pricing pages: Intelligence &
|
|
18
18
|
# Coding indexes and 7:2:1 blended prices all match (AA field price1mBlended7To2To1);
|
|
19
19
|
# coding cost-per-task is chart-only (not independently reconstructed). Prices are
|
|
20
20
|
# standard short-context API tier — on subscription CLIs treat $ as relative ranking.
|
|
21
21
|
# Your own job outcomes (~/.omnilane/jobs/) outrank these priors; edit lanes to match.
|
|
22
22
|
|
|
23
|
-
hardest-coding:
|
|
24
|
-
bulk-mechanical: codex gpt-5.6-
|
|
25
|
-
triage:
|
|
26
|
-
hard-judgment:
|
|
27
|
-
taste-final:
|
|
28
|
-
consult:
|
|
29
|
-
ui-draft:
|
|
30
|
-
long-context:
|
|
31
|
-
fast-agentic:
|
|
32
|
-
live-search:
|
|
33
|
-
coding-overflow: grok grok-4.
|
|
34
|
-
arbitrate:
|
|
23
|
+
hardest-coding: claude claude-fable-5-1 xhigh | codex gpt-5.6-sol xhigh # ordered on coding capability: the leading Claude tier wins both coding components; Sol remains the established Codex-harness value fallback
|
|
24
|
+
bulk-mechanical: codex gpt-5.6-sol high | gemini "Gemini 3.7 Flash (High)" - | claude claude-sonnet-5 high # ordered on endurance per dollar: Sol dominates Terra within Codex; current Flash is the faster, cheaper middle fallback; Sonnet preserves subscription quota
|
|
25
|
+
triage: codex gpt-5.6-luna high | gemini "Gemini 3.7 Flash (Low)" - | claude claude-haiku-4-5 - # ordered on cost per task at usable intelligence: Luna high buys a meaningful quality lift cheaply; Flash and Haiku are low-cost cross-vendor fallbacks
|
|
26
|
+
hard-judgment: claude claude-fable-5-1 xhigh | codex gpt-5.6-sol max | grok grok-4.6 - # ordered on agentic knowledge work: Fable leads the Claude field; Sol stays ahead of Grok because Grok effort is ignored and its reproduced row is unknown
|
|
27
|
+
taste-final: claude claude-fable-5-1 high | codex gpt-5.6-sol max # ordered on prose and polish: Fable leads Opus on intelligence and factual breadth; Sol is the cross-vendor fallback
|
|
28
|
+
consult: codex gpt-5.6-sol max | claude claude-fable-5-1 high | grok grok-4.6 - | gemini "Gemini 3.7 Flash (High)" - # direct named-model chain uses the strongest current Claude and Flash slots; keep --vendor to prevent fallback
|
|
29
|
+
ui-draft: codex gpt-5.6-sol xhigh | claude claude-fable-5-1 high # ordered for drafts with a design system or reference images: Sol leads measured multimodal and coding evidence; Fable follows for polish
|
|
30
|
+
long-context: gemini "Gemini 3.7 Flash (Medium)" - | codex gpt-5.6-terra max | claude claude-opus-5 medium # ordered on long-context reasoning, then cost and throughput: Flash leads; Terra matches its long-context result; Opus is the cheaper Claude fallback
|
|
31
|
+
fast-agentic: gemini "Gemini 3.7 Flash (Medium)" - | codex gpt-5.6-luna high # ordered on interactive tool-loop latency: Flash gives up little agentic quality for far faster first output; Luna high is the low-latency Codex fallback
|
|
32
|
+
live-search: grok grok-4.6 - | off # native X and web search lane; no real substitute
|
|
33
|
+
coding-overflow: grok grok-4.6 - | gemini "Gemini 3.7 Flash (High)" - | kimi kimi-k3 - | qwen qwen3-coder-plus - | opencode - - | off # coding relief ordered by capability and value: Grok has the lowest frontier hallucination rate; Flash is the cheapest strong coder here; revisit the best-value Qwen tier when its CLI alias can be verified
|
|
34
|
+
arbitrate: off - - # opinion panel remains opt-in because each voter and round consumes quota
|
|
35
35
|
# Enable: `arbitrate: vote codex,claude,grok -` (any 1-4 of codex/claude/grok/gemini)
|
|
36
36
|
# Debate round (each voter rebuts the others): set the effort field to 2.
|
|
37
37
|
# Custom gate: `arbitrate: exec /path/to/script -`
|
|
38
|
-
# Claude Fable 5
|
|
39
|
-
#
|
|
40
|
-
#
|
|
41
|
-
#
|
|
42
|
-
#
|
|
43
|
-
#
|
|
44
|
-
#
|
|
45
|
-
# taste-final: claude claude-fable-5 high
|
|
46
|
-
# in ~/.omnilane/routing.local.yaml.
|
|
38
|
+
# Claude Fable 5.1 is in the judgment, taste, and hardest-coding defaults because
|
|
39
|
+
# it leads Opus 5 on every Artificial Analysis axis at the same effort.
|
|
40
|
+
# It is not in bulk or triage: it prices at twice Opus 5 per token and consumes
|
|
41
|
+
# the most subscription quota per turn. Opus 5 remains the lower-hallucination,
|
|
42
|
+
# lower-price Claude choice and can return to any lane via
|
|
43
|
+
# ~/.omnilane/routing.local.yaml, for example:
|
|
44
|
+
# hard-judgment: claude claude-opus-5 xhigh
|
package/scripts/configure.sh
CHANGED
|
@@ -151,10 +151,10 @@ esac
|
|
|
151
151
|
# Dynamic/API catalogs stay curated — "c" always accepts an exact model ID.
|
|
152
152
|
CODEX_MODELS=("gpt-5.6" "gpt-5.6-sol" "gpt-5.6-terra" "gpt-5.6-luna" "gpt-5.5" "gpt-5.4" "gpt-5.4-mini" "gpt-5.3-codex-spark")
|
|
153
153
|
CODEX_EFFORTS=("xhigh" "max" "ultra" "high" "medium" "low" "minimal" "none")
|
|
154
|
-
CLAUDE_MODELS=("default" "best" "fable" "opus" "sonnet" "haiku" "opus[1m]" "sonnet[1m]" "opusplan" "claude-fable-5" "claude-opus-5" "claude-sonnet-5" "claude-opus-4-8" "claude-opus-4-7" "claude-opus-4-6" "claude-opus-4-5-20251101" "claude-sonnet-4-6" "claude-sonnet-4-5-20250929" "claude-haiku-4-5" "claude-haiku-4-5-20251001")
|
|
154
|
+
CLAUDE_MODELS=("default" "best" "fable" "opus" "sonnet" "haiku" "opus[1m]" "sonnet[1m]" "opusplan" "claude-fable-5" "claude-fable-5-1" "claude-opus-5" "claude-sonnet-5" "claude-opus-4-8" "claude-opus-4-7" "claude-opus-4-6" "claude-opus-4-5-20251101" "claude-sonnet-4-6" "claude-sonnet-4-5-20250929" "claude-haiku-4-5" "claude-haiku-4-5-20251001")
|
|
155
155
|
CLAUDE_EFFORTS=("max" "xhigh" "high" "medium" "low" "-")
|
|
156
|
-
GEMINI_MODELS=("gemini-3.
|
|
157
|
-
GROK_MODELS=("grok-4.
|
|
156
|
+
GEMINI_MODELS=("gemini-3.7-flash-high" "gemini-3.7-flash-medium" "gemini-3.7-flash-low" "gemini-3.6-flash-high" "gemini-3.6-flash-medium" "gemini-3.6-flash-low" "gemini-3.1-pro-high" "gemini-3.1-pro-low" "claude-sonnet-4-6" "claude-opus-4-6-thinking" "gpt-oss-120b-medium")
|
|
157
|
+
GROK_MODELS=("grok-4.6" "headroom-grok-build" "grok-4.3-official")
|
|
158
158
|
KIMI_MODELS=("kimi-k3" "kimi-k2.7-code" "kimi-k2.5")
|
|
159
159
|
QWEN_MODELS=("qwen3.7-max" "qwen3.7-plus" "qwen3.6-plus" "qwen3.5-plus" "qwen3-max-2026-01-23" "qwen3-coder-next" "qwen3-coder-plus" "qwen3-coder-flash")
|
|
160
160
|
# OpenCode models use provider/model form; OpenRouter models use catalog slugs.
|
|
@@ -31,9 +31,9 @@ trap cleanup_temp_files EXIT
|
|
|
31
31
|
voter_spec() { # vendor -> "model<TAB>effort"
|
|
32
32
|
case "$1" in
|
|
33
33
|
codex) printf 'gpt-5.6-sol\thigh' ;;
|
|
34
|
-
claude) printf 'claude-
|
|
35
|
-
gemini) printf 'Gemini 3.
|
|
36
|
-
grok) printf 'grok-4.
|
|
34
|
+
claude) printf 'claude-fable-5-1\thigh' ;;
|
|
35
|
+
gemini) printf 'Gemini 3.7 Flash (High)\t-' ;;
|
|
36
|
+
grok) printf 'grok-4.6\t-' ;;
|
|
37
37
|
*) return 1 ;;
|
|
38
38
|
esac
|
|
39
39
|
}
|
package/skills/omnilane/SKILL.md
CHANGED
|
@@ -50,28 +50,26 @@ what dispatch picks when the first-choice vendor CLI is not installed.
|
|
|
50
50
|
|
|
51
51
|
| Lane | First choice | Backup | When |
|
|
52
52
|
|---|---|---|---|
|
|
53
|
-
| hardest-coding |
|
|
54
|
-
| bulk-mechanical | GPT-5.6
|
|
55
|
-
| triage | GPT-5.6 Luna (
|
|
56
|
-
| hard-judgment | Claude
|
|
57
|
-
| taste-final | Claude
|
|
58
|
-
| consult |
|
|
59
|
-
| ui-draft | GPT-5.6 Sol (xhigh) | Claude
|
|
60
|
-
| long-context | Gemini 3.
|
|
61
|
-
| fast-agentic |
|
|
62
|
-
| live-search | Grok 4.
|
|
63
|
-
| coding-overflow | Grok 4.
|
|
53
|
+
| hardest-coding | Claude Fable 5.1 (xhigh) | GPT-5.6 Sol (xhigh) | Hardest implementation, deep root-cause debug, correctness-critical edits |
|
|
54
|
+
| bulk-mechanical | GPT-5.6 Sol (high) | Gemini 3.7 Flash (High) → Claude Sonnet 5 (high) | Refactors, migrations, tests, review sweeps — mechanical endurance |
|
|
55
|
+
| triage | GPT-5.6 Luna (high) | Gemini 3.7 Flash (Low) → Claude Haiku 4.5 | High-volume scans, first-pass filtering |
|
|
56
|
+
| hard-judgment | Claude Fable 5.1 (xhigh) | GPT-5.6 Sol (max) → Grok 4.6 | Architecture arbitration, deep reasoning, second opinions |
|
|
57
|
+
| taste-final | Claude Fable 5.1 (high) | GPT-5.6 Sol (max) | User-facing prose, prompt/doc polish, Chinese phrasing, style arbitration |
|
|
58
|
+
| consult | GPT-5.6 Sol (max) | Claude Fable 5.1 (high) → Grok 4.6 → Gemini 3.7 Flash (High) | Direct named-model consultation; always keep `--vendor` |
|
|
59
|
+
| ui-draft | GPT-5.6 Sol (xhigh) | Claude Fable 5.1 (high) | UI drafts only WITH a design system / reference images; open-ended visual taste goes to taste-final |
|
|
60
|
+
| long-context | Gemini 3.7 Flash (Medium) | GPT-5.6 Terra (max) → Claude Opus 5 (medium) | Long-context synthesis ordered on AA-LCR, then cost and throughput |
|
|
61
|
+
| fast-agentic | Gemini 3.7 Flash (Medium) | GPT-5.6 Luna (high) | Fast multi-step agentic loops, multimodal checks |
|
|
62
|
+
| live-search | Grok 4.6 | — (off) | Realtime X/web search and social context |
|
|
63
|
+
| coding-overflow | Grok 4.6 | Gemini 3.7 Flash (High) → Kimi K3 → Qwen3 Coder Plus → OpenCode | Codex-quota relief valve for mid-tier coding; verify factual claims |
|
|
64
64
|
| arbitrate | off (opt-in vote panel) | — | Disabled by default. Enable with `arbitrate: vote codex,claude,grok -` in routing.local.yaml or via the configurator (any 1-4 voters). One quota hit PER VOTER PER ROUND; you chair: read the opinions and own the decision. Effort field 2 = debate round (voters rebut each other) |
|
|
65
65
|
|
|
66
|
-
Claude Fable 5 (`claude-fable-5`) is
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
it anyway, select it in the configurator or override a lane in
|
|
74
|
-
`~/.omnilane/routing.local.yaml` (e.g. `taste-final: claude claude-fable-5 high`).
|
|
66
|
+
Claude Fable 5.1 (`claude-fable-5-1`) is in the judgment, taste, and
|
|
67
|
+
hardest-coding defaults because it leads Opus 5 on every Artificial Analysis
|
|
68
|
+
axis at the same effort. It is not in bulk or triage because it prices at twice
|
|
69
|
+
Opus 5 per token and consumes the most subscription quota per turn. Opus 5
|
|
70
|
+
remains the lower-hallucination, lower-price Claude choice and can return to any
|
|
71
|
+
lane via `~/.omnilane/routing.local.yaml`, for example:
|
|
72
|
+
`hard-judgment: claude claude-opus-5 xhigh`.
|
|
75
73
|
|
|
76
74
|
## Natural-language consultation
|
|
77
75
|
|
|
@@ -91,15 +89,15 @@ Users may speak normally; they do not need lane names.
|
|
|
91
89
|
| Alias | Vendor | Model | Effort |
|
|
92
90
|
|---|---|---|---|
|
|
93
91
|
| Opus | claude | claude-opus-5 | high |
|
|
94
|
-
| Fable | claude | claude-fable-5 | high |
|
|
92
|
+
| Fable 5.1 | claude | claude-fable-5-1 | high |
|
|
95
93
|
| Sonnet | claude | claude-sonnet-5 | high |
|
|
96
94
|
| Haiku | claude | claude-haiku-4-5 | - |
|
|
97
95
|
| Sol | codex | gpt-5.6-sol | max |
|
|
98
96
|
| Terra | codex | gpt-5.6-terra | max |
|
|
99
|
-
| Luna | codex | gpt-5.6-luna |
|
|
100
|
-
| Grok 4.
|
|
101
|
-
| Gemini Pro | gemini | Gemini 3.1 Pro (High) | - |
|
|
102
|
-
| Gemini Flash | gemini | Gemini 3.
|
|
97
|
+
| Luna | codex | gpt-5.6-luna | high |
|
|
98
|
+
| Grok 4.6 | grok | grok-4.6 | - |
|
|
99
|
+
| Gemini 3.1 Pro | gemini | Gemini 3.1 Pro (High) | - |
|
|
100
|
+
| Gemini 3.7 Flash | gemini | Gemini 3.7 Flash (High) | - |
|
|
103
101
|
| Kimi | kimi | kimi-k3 | - |
|
|
104
102
|
| Qwen | qwen | qwen3-coder-plus | - |
|
|
105
103
|
| OpenCode | opencode | provider/model form, or `-` for its own default | - |
|
|
@@ -159,19 +157,24 @@ dispatch stay in this skill and the CLI. Manage the local board with
|
|
|
159
157
|
|
|
160
158
|
## Per-model notes (apply the row matching YOUR main model)
|
|
161
159
|
|
|
162
|
-
- **Claude
|
|
163
|
-
|
|
164
|
-
|
|
160
|
+
- **Claude Fable 5.1 main**: hard judgment, taste finalization, and the hardest
|
|
161
|
+
coding are yours. Dispatch bulk work to Sol high and long-context or fast
|
|
162
|
+
loops to Gemini 3.7 Flash.
|
|
163
|
+
- **Claude Opus 5 main**: judgment and taste remain strong self-execute lanes;
|
|
164
|
+
use local overrides when its lower hallucination rate or price is preferred.
|
|
165
165
|
- **Claude Sonnet main**: coordination/tools/mid-tier coding only; never
|
|
166
166
|
self-assign top judgment or hardest implementation.
|
|
167
167
|
- **GPT Sol main**: hardest coding + hard judgment are yours (use max for
|
|
168
168
|
judgment turns, xhigh for coding); cross to taste-final for style calls.
|
|
169
|
-
- **GPT Terra main**:
|
|
170
|
-
|
|
171
|
-
|
|
172
|
-
|
|
173
|
-
|
|
174
|
-
|
|
175
|
-
- **Gemini 3.
|
|
176
|
-
are yours
|
|
177
|
-
|
|
169
|
+
- **GPT Terra main**: long-context Codex fallback work is yours at max;
|
|
170
|
+
bulk-mechanical now defaults to Sol high, and genuinely hardest pieces
|
|
171
|
+
escalate to Sol xhigh.
|
|
172
|
+
- **Grok 4.6 main**: live-search and coding overflow are yours; its measured
|
|
173
|
+
hallucination rate is the lowest among the frontier rows, but still verify
|
|
174
|
+
every API signature and cited fact before shipping.
|
|
175
|
+
- **Gemini 3.7 Flash main**: long-context and fast agentic/multimodal loops
|
|
176
|
+
are yours at the lane's configured effort; bulk and overflow use the high row.
|
|
177
|
+
Never self-assign top judgment.
|
|
178
|
+
- **Gemini 3.1 Pro main**: it remains directly selectable, but the default
|
|
179
|
+
long-context lane now prefers Gemini 3.7 Flash on LCR, cost, and throughput;
|
|
180
|
+
route hardest coding and judgment to the stronger Codex and Claude lanes.
|