omnilane 0.15.0 → 0.21.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,172 @@
1
+ ---
2
+ name: omnilane
3
+ description: 'Universal model-routing table + cross-vendor dispatch for ANY harness (Claude Code, Codex, Grok Build, Antigravity). Use when delegating subtasks, choosing a model for work, planning multi-part tasks, or when asked about model routing, delegate, dispatch, which model, tier selection, escalate, 派工, 模型路由. One routing table; the main loop self-executes its own lane and shells out to every other vendor via dispatch.sh.'
4
+ ---
5
+
6
+ # omnilane — one routing table, every harness
7
+
8
+ You (the main loop) may be Claude, GPT, Grok, or Gemini. The procedure is identical:
9
+
10
+ 1. **Identify your main model.** You know which model you are running as.
11
+ 2. **Split the work into subtasks and classify each into a lane** (table below).
12
+ 3. **If the lane's model is you, self-execute.** Otherwise dispatch:
13
+ `<repo>/scripts/dispatch.sh [--vendor V] [--mode work] [--workdir DIR] <lane> "<task>"`
14
+ Add `--background` for long tasks; poll with `scripts/jobs.sh status|result <id>`.
15
+ Before changing lane order from anecdotal outcomes, run
16
+ `scripts/jobs.sh recommend [--last N] [--lane L] [--min-samples N]` and report
17
+ its evidence threshold. The command is read-only and never changes routing.
18
+ Preview old completed-job cleanup with `scripts/jobs.sh prune --keep <N>`;
19
+ deletion requires the explicit `--apply` flag and never targets running jobs.
20
+ A deep task whose CLI call may outrun the 600s per-call watchdog can raise its
21
+ cap with `--timeout <seconds>` (e.g. `--timeout 1200` for hard-judgment /
22
+ long-context). It bounds each CLI call, not the whole dispatch.
23
+ For one aggregate fuse across lock wait, retries, voters, and rounds, add
24
+ `--job-timeout <seconds>`. It is disabled by default; deep full-repository
25
+ audits typically need 7200–14400 seconds, and expiry returns 124. The one
26
+ automatic exception is non-Git Codex `work`: without an explicit, lane, or
27
+ global job timeout, its resolved per-call timeout becomes the whole-job fuse,
28
+ capped at the supervisor's 999999999-second maximum. If the bundled Perl
29
+ supervisor is unavailable, it warns and continues through the existing
30
+ per-call watchdog path.
31
+
32
+ Run `scripts/dispatch.sh --list` to see the effective table (local overrides win).
33
+ When routing is unexpectedly unavailable, run `bin/omnilane doctor` before
34
+ changing configuration; it reports state and dependencies without repairing them.
35
+ Doctor remains offline unless the operator explicitly adds `--probe V`; that
36
+ bounded probe returns metadata only. Use `bin/omnilane benchmark` for a fixed
37
+ no-call route plan, and add `--run` only when actual advise-mode comparison calls
38
+ were explicitly requested. Neither command changes routing.
39
+ Lanes are fallback chains — dispatch uses the first vendor CLI actually installed,
40
+ so the same table works with any subset of subscriptions.
41
+
42
+ ## Lanes (defaults; see routing.yaml for the live values)
43
+
44
+ Each lane's **backup** is the next candidate in its `routing.yaml` chain —
45
+ what dispatch picks when the first-choice vendor CLI is not installed.
46
+
47
+ | Lane | First choice | Backup | When |
48
+ |---|---|---|---|
49
+ | hardest-coding | GPT-5.6 Sol (xhigh) | Claude Opus 5 (xhigh) | Hardest implementation, deep root-cause debug, correctness-critical edits |
50
+ | bulk-mechanical | GPT-5.6 Terra (max) | Claude Sonnet 5 (high) | Refactors, migrations, tests, review sweeps — mechanical endurance |
51
+ | triage | GPT-5.6 Luna (medium) | Gemini 3.6 Flash (Low) | High-volume scans, first-pass filtering |
52
+ | hard-judgment | Claude Opus 5 (xhigh) | GPT-5.6 Sol (max) | Architecture arbitration, deep reasoning, second opinions |
53
+ | taste-final | Claude Opus 5 (high) | GPT-5.6 Sol (max) | User-facing prose, prompt/doc polish, Chinese phrasing, style arbitration |
54
+ | consult | Explicit named vendor/model | — (no fallback) | Direct natural-language consultation; always keep `--vendor` |
55
+ | ui-draft | GPT-5.6 Sol (xhigh) | Claude Opus 5 (high) | UI drafts only WITH a design system / reference images; open-ended visual taste goes to taste-final |
56
+ | long-context | Gemini 3.1 Pro (High) | GPT-5.6 Sol (high) | 1M-token synthesis; Pro is agentic-capable, while fast repeated loops prefer Flash on speed/cost |
57
+ | fast-agentic | GPT-5.6 Luna (max) | Gemini 3.6 Flash (High) | Fast multi-step agentic loops, multimodal checks |
58
+ | live-search | Grok 4.5 | — (off) | Realtime X/web search and social context |
59
+ | coding-overflow | Grok 4.5 | Kimi K3 → Qwen3 Coder Plus → OpenCode | Codex-quota relief valve for mid-tier coding; verify factual claims |
60
+ | arbitrate | off (opt-in vote panel) | — | Disabled by default. Enable with `arbitrate: vote codex,claude,grok -` in routing.local.yaml or via the configurator (any 1-4 voters). One quota hit PER VOTER PER ROUND; you chair: read the opinions and own the decision. Effort field 2 = debate round (voters rebut each other) |
61
+
62
+ Claude Fable 5 (`claude-fable-5`) is absent from the defaults on purpose: the
63
+ top Claude tier is usually the main loop itself, not a dispatched worker, and
64
+ it prices at twice Opus 5. This is a cost / guardrail / main-loop policy choice,
65
+ not a capability verdict — Artificial Analysis calls Opus 5 (61) and Fable 5 (60)
66
+ "effectively tied" on the Intelligence Index, but Opus 5 leads AA-Briefcase by
67
+ 146 Elo at 20% lower cost per task. Fable 5 keeps the lead on factual breadth
68
+ (AA-Omniscience), so name it explicitly for recall-heavy consults. To route to
69
+ it anyway, select it in the configurator or override a lane in
70
+ `~/.omnilane/routing.local.yaml` (e.g. `taste-final: claude claude-fable-5 high`).
71
+
72
+ ## Natural-language consultation
73
+
74
+ Users may speak normally; they do not need lane names.
75
+
76
+ 1. Capability-only question (`which model`, `what can Claude do`, `哪個模型`,
77
+ `誰適合`) → classify the need, then answer with the first available model
78
+ shown for that lane by `dispatch.sh --list`; do not dispatch unless execution
79
+ is also requested.
80
+ 2. Generic vendor name (`Claude`, `Codex`, `Grok`, `Gemini`, `OpenCode`) → run
81
+ `dispatch.sh --vendor <vendor> consult "<task>"`.
82
+ 3. Canonical model alias → pass its vendor, model, and effort from the table
83
+ below. Never silently substitute another model family.
84
+ 4. No named target → classify into an existing lane and dispatch normally.
85
+ 5. Unknown or ambiguous nickname → ask for clarification; do not guess or run.
86
+
87
+ | Alias | Vendor | Model | Effort |
88
+ |---|---|---|---|
89
+ | Opus | claude | claude-opus-5 | high |
90
+ | Fable | claude | claude-fable-5 | high |
91
+ | Sonnet | claude | claude-sonnet-5 | high |
92
+ | Haiku | claude | claude-haiku-4-5 | - |
93
+ | Sol | codex | gpt-5.6-sol | max |
94
+ | Terra | codex | gpt-5.6-terra | max |
95
+ | Luna | codex | gpt-5.6-luna | medium |
96
+ | Grok 4.5 | grok | grok-4.5 | - |
97
+ | Gemini Pro | gemini | Gemini 3.1 Pro (High) | - |
98
+ | Gemini Flash | gemini | Gemini 3.6 Flash (High) | - |
99
+ | Kimi | kimi | kimi-k3 | - |
100
+ | Qwen | qwen | qwen3-coder-plus | - |
101
+ | OpenCode | opencode | provider/model form, or `-` for its own default | - |
102
+ | OpenRouter | openrouter | explicit OpenRouter slug (e.g. anthropic/claude-sonnet-5) | - |
103
+
104
+ OpenCode is the multi-provider aggregator CLI (75+ providers): work-capable,
105
+ last resort in coding-overflow. OpenRouter is direct-API — no CLI needed, only
106
+ `OPENROUTER_API_KEY` — and is **advise/consult only** (it cannot edit files);
107
+ its model slug is mandatory. "Ask <any hosted model> via OpenRouter" →
108
+ `dispatch.sh --vendor openrouter --model <slug> consult "<task>"`.
109
+
110
+ Examples:
111
+
112
+ - Ask Opus to challenge this architecture →
113
+ `dispatch.sh --vendor claude --model claude-opus-5 --effort high consult "challenge this architecture"`
114
+ - 請 Grok 查最新公開資訊 →
115
+ `dispatch.sh --vendor grok consult "查最新公開資訊"`
116
+ - 哪個模型適合檢查大型 repo? → answer only; do not dispatch.
117
+
118
+ Consultation defaults to `advise`. Use `--mode work --workdir <dir>` only for
119
+ an explicit edit request. Missing explicit targets fail clearly; never remove
120
+ `--vendor` to obtain a fallback.
121
+
122
+ ## Live UI is observation only
123
+
124
+ The optional Live UI is a read-only observer, not a prompt or dispatch path.
125
+ It displays existing jobs' `task.txt` and public `out.txt`, but never raw logs;
126
+ its history search and state filters can export only the currently visible public
127
+ metadata as local JSON; tokens and task/result bodies are excluded from export.
128
+ it cannot interpret natural language, choose routes, dispatch, retry, cancel,
129
+ delete jobs, or edit configuration. Natural-language interpretation and
130
+ dispatch stay in this skill and the CLI. Manage the local board with
131
+ `omnilane ui start|status|url|stop`, and stop it when monitoring is finished.
132
+
133
+ ## Rules
134
+
135
+ - **Dispatch in `advise` mode by default** (read-only worker). Use `--mode work`
136
+ only when the worker must edit files, and give it an explicit `--workdir`.
137
+ - **`--mode sysops`** is `work` minus the vendor sandbox, for service
138
+ operations the sandbox denies (launchctl, system daemons). Codex runs with
139
+ `-s danger-full-access`; other vendors treat it as `work`. Explicit
140
+ per-dispatch opt-in only — never a lane default, and the task text must
141
+ name the exact service commands the worker is authorized to run.
142
+ Codex `work`/`sysops` still needs a git-repo `--workdir` (non-git
143
+ directories trip the whole-job fuse).
144
+ - **Every dispatched task states acceptance criteria and the exact verification
145
+ command.** Do not accept "done" without evidence.
146
+ - **No nested dispatch**: workers must not fan out again (enforced via
147
+ `OMNILANE_DEPTH`). Escalate back to the main loop instead.
148
+ - **Same-directory codex dispatches are serialized automatically** (lock);
149
+ do not try to parallelize them yourself.
150
+ - Escalate without asking: two failed attempts on a lane → move one lane up
151
+ (triage → bulk-mechanical → hardest-coding).
152
+ - Vendor quota exhausted (429 / "stream disconnected" / usage-limit message):
153
+ send mid-tier coding through coding-overflow instead; never silently downgrade
154
+ hardest-coding — wait or escalate to the user.
155
+
156
+ ## Per-model notes (apply the row matching YOUR main model)
157
+
158
+ - **Claude (Fable/Opus main)**: top judgment and taste are yours — self-execute;
159
+ push mechanical coding volume out to the codex lanes.
160
+ - **Claude Sonnet main**: coordination/tools/mid-tier coding only; never
161
+ self-assign top judgment or hardest implementation.
162
+ - **GPT Sol main**: hardest coding + hard judgment are yours (use max for
163
+ judgment turns, xhigh for coding); cross to taste-final for style calls.
164
+ - **GPT Terra main**: bulk work is yours at max; escalate the genuinely hardest
165
+ pieces to Sol instead of grinding.
166
+ - **Grok 4.5 main**: mid-tier coding + live-search are yours; verify every API
167
+ signature and cited fact before shipping (measured high hallucination rate).
168
+ - **Gemini Flash main**: fast agentic/multimodal loops are yours; never
169
+ self-assign top judgment.
170
+ - **Gemini 3.1 Pro main**: 1M-context synthesis and context-heavy agentic work
171
+ are yours. Prefer Gemini Flash for fast repeated tool loops on speed/cost;
172
+ route hardest coding and judgment to the stronger codex lanes.