omnilane 0.42.9 → 0.45.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +2 -2
- package/.claude-plugin/plugin.json +1 -1
- package/CHANGELOG.md +260 -1
- package/README.ja.md +190 -32
- package/README.ko.md +190 -32
- package/README.md +319 -73
- package/README.zh-CN.md +190 -31
- package/README.zh-TW.md +180 -53
- package/VERSION +1 -1
- package/bin/omnilane +29 -0
- package/completions/_omnilane +1 -1
- package/completions/omnilane.bash +1 -1
- package/completions/omnilane.fish +2 -0
- package/config/aa-model-policy.json +1339 -1170
- package/docs/aa-model-coverage-2026-09-05.json +43 -1
- package/docs/model-capabilities-2026-09.md +195 -4
- package/docs/native-executor.md +59 -0
- package/package.json +1 -1
- package/plugin.json +1 -1
- package/routing.yaml +25 -23
- package/scripts/aa_rebaseline.py +409 -0
- package/scripts/configure.sh +1 -1
- package/scripts/dispatch.sh +54 -4
- package/scripts/doctor.sh +7 -1
- package/scripts/lib/aa_lanes.py +135 -0
- package/scripts/lib/aa_policy.py +165 -8
- package/scripts/lib/build_overlay.py +54 -18
- package/scripts/lib/caller_identity.py +75 -13
- package/scripts/lib/cli_provenance.py +82 -0
- package/scripts/lib/native.py +107 -2
- package/scripts/lib/native_context.py +145 -0
- package/scripts/lib/overlay_health.py +24 -2
- package/scripts/lib/probe.py +2 -0
- package/scripts/lib/probe_sweep.py +189 -0
- package/scripts/lib/resign.py +493 -0
- package/scripts/release-audit.sh +31 -0
- package/skills/omnilane/SKILL.md +393 -477
package/README.md
CHANGED
|
@@ -44,42 +44,170 @@ CLI or seven — dispatch picks the first candidate you actually have, and a lan
|
|
|
44
44
|
with nothing available simply turns off. The default table works on a single
|
|
45
45
|
subscription.
|
|
46
46
|
|
|
47
|
-
**[⬇
|
|
47
|
+
**[⬇ 60-second start](#-60-second-start)** · **[🤖 Let your AI assistant drive it](#-let-your-ai-assistant-drive-omnilane)** · **[❓ FAQ](#-faq)**
|
|
48
48
|
|
|
49
49
|
## ⚡ 60-second start
|
|
50
50
|
|
|
51
|
-
|
|
51
|
+
You, a person at a terminal, can dispatch right away.
|
|
52
|
+
|
|
53
|
+
**1. Install.**
|
|
52
54
|
|
|
53
55
|
```bash
|
|
54
|
-
npm i -g omnilane
|
|
55
|
-
export OMNILANE_AA_OPERATOR_ASSERTED_HUMAN=1 # you are the operator, not a model
|
|
56
|
-
omnilane route hardest-coding "fix the flaky auth token refresh"
|
|
57
|
-
omnilane doctor # see which AI CLIs / keys you have
|
|
58
|
-
omnilane ui start # optional: watch jobs live in your browser
|
|
56
|
+
npm i -g omnilane
|
|
59
57
|
```
|
|
60
58
|
|
|
61
|
-
|
|
59
|
+
Or clone it, which also gives you the routing table and the skill to customise:
|
|
62
60
|
|
|
63
61
|
```bash
|
|
64
62
|
git clone https://github.com/Seraphim0916/omnilane && cd omnilane
|
|
65
63
|
./install.sh # finds your CLIs, links the skill, speaks your language
|
|
66
|
-
|
|
64
|
+
```
|
|
65
|
+
|
|
66
|
+
**2. See what you have.** `doctor` lists which model CLIs and API keys omnilane
|
|
67
|
+
can reach, so you know what will actually run. It changes nothing.
|
|
68
|
+
|
|
69
|
+
```bash
|
|
70
|
+
omnilane doctor
|
|
71
|
+
omnilane list # the routing table this machine resolves
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
**3. Say you are the operator, then dispatch.**
|
|
75
|
+
|
|
76
|
+
```bash
|
|
77
|
+
export OMNILANE_AA_OPERATOR_ASSERTED_HUMAN=1
|
|
67
78
|
omnilane route hardest-coding "fix the flaky auth token refresh"
|
|
79
|
+
omnilane ui start # optional: watch jobs live in your browser
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
> **Why the export?** omnilane checks every delegation against the capability
|
|
83
|
+
> score of whoever is asking, so a dispatch has to say who that is. A human says
|
|
84
|
+
> it once with `OMNILANE_AA_OPERATOR_ASSERTED_HUMAN=1` (or `--operator-asserted-human`
|
|
85
|
+
> per call). A model cannot say it for itself: its identity is read from the CLI
|
|
86
|
+
> that launched it. With neither, the dispatch is refused with
|
|
87
|
+
> `missing-caller-context` before any job exists.
|
|
88
|
+
|
|
89
|
+
That is all a human needs. The rest of this section is for the more useful
|
|
90
|
+
setup: your AI assistant dispatching on its own.
|
|
91
|
+
|
|
92
|
+
## 🤖 Let your AI assistant drive omnilane
|
|
93
|
+
|
|
94
|
+
The assistant (Claude Code, Codex, Grok Build or Antigravity) reads a skill file
|
|
95
|
+
that tells it how to pick a lane and dispatch. Four steps, once per machine.
|
|
96
|
+
|
|
97
|
+
### Step 1 — Give the assistant the skill
|
|
98
|
+
|
|
99
|
+
`./install.sh` links it for every CLI it finds. By hand:
|
|
100
|
+
|
|
101
|
+
| Assistant | How |
|
|
102
|
+
|---|---|
|
|
103
|
+
| Claude Code | `claude plugin marketplace add <this repo>` then `claude plugin install omnilane@omnilane` (also gives `/route`, `/route-jobs` and the completion inbox), or link `skills/omnilane` into `~/.claude/skills/` |
|
|
104
|
+
| Codex | link `skills/omnilane` into `~/.codex/skills/` |
|
|
105
|
+
| Grok Build | `grok plugin install <this repo> --trust` |
|
|
106
|
+
| Antigravity | `agy plugin install <this repo>` (check first with `agy plugin validate <this repo>`) |
|
|
107
|
+
|
|
108
|
+
### Step 2 — Prove, once, that each CLI selects the model it is asked for
|
|
109
|
+
|
|
110
|
+
A model caller is only allowed to dispatch to a target this machine has *proven*:
|
|
111
|
+
that `codex -m gpt-5.6-sol` really runs Sol, and so on. The proof is a local file,
|
|
112
|
+
the **transport overlay**. Nothing ships with one. Without it every lane refuses
|
|
113
|
+
a model caller with `runtime-mapping-unverified`, and `omnilane doctor` warns
|
|
114
|
+
`no overlay configured`.
|
|
115
|
+
|
|
116
|
+
Build it from a normal desktop terminal. (An ssh login cannot read the keychain
|
|
117
|
+
the CLIs log in with, so it would report them all as not logged in.)
|
|
118
|
+
|
|
119
|
+
```bash
|
|
120
|
+
cd "$(npm root -g)/omnilane" # or your clone
|
|
121
|
+
ROOT=~/.omnilane/transport-evidence/first-sweep
|
|
122
|
+
python3 scripts/lib/probe_sweep.py --root "$ROOT" # one tiny prompt per selector, about 55 calls
|
|
123
|
+
python3 scripts/lib/build_overlay.py --root "$ROOT"
|
|
124
|
+
cp "$ROOT/transport-contracts.local.json" ~/.omnilane/transport-contracts.local.json
|
|
125
|
+
echo 'export OMNILANE_AA_TRANSPORT_OVERLAY="$HOME/.omnilane/transport-contracts.local.json"' >> ~/.omnilane/local.sh
|
|
126
|
+
omnilane doctor | grep transport-overlay # PASS, with a count per vendor
|
|
127
|
+
```
|
|
128
|
+
|
|
129
|
+
A vendor you are not logged in to is reported `unprobeable` and simply stays
|
|
130
|
+
unverified; the others work.
|
|
131
|
+
|
|
132
|
+
### Step 3 — Keep the proof current without doing it by hand
|
|
133
|
+
|
|
134
|
+
The overlay pins each CLI executable by hash, and **the CLIs update themselves**,
|
|
135
|
+
often weekly. After an update that vendor's lanes are refused until the overlay
|
|
136
|
+
is re-signed. `omnilane resign` does the whole job: finds what changed, re-probes
|
|
137
|
+
only that vendor, checks the result, swaps it in, sends one real dispatch to
|
|
138
|
+
confirm, and restores the old file if that fails.
|
|
139
|
+
|
|
140
|
+
It will not re-sign just anything. A changed CLI is re-signed **unattended** only
|
|
141
|
+
when it carries the same code-signing team as the one on record and sits in the
|
|
142
|
+
same install location. So tell it once which signers you accept:
|
|
143
|
+
|
|
144
|
+
```bash
|
|
145
|
+
omnilane resign --record-signers # once, right after Step 2
|
|
146
|
+
```
|
|
147
|
+
|
|
148
|
+
If you patch a vendor CLI yourself after every update and re-sign it adhoc,
|
|
149
|
+
tell omnilane that too, once per vendor; an adhoc update in the same install
|
|
150
|
+
directory is then re-signed unattended as well:
|
|
151
|
+
|
|
152
|
+
```bash
|
|
153
|
+
omnilane resign --trust-adhoc claude # only if you re-sign claude adhoc yourself
|
|
154
|
+
```
|
|
155
|
+
|
|
156
|
+
Then let it run every day. Any scheduler works as long as it runs **inside your
|
|
157
|
+
desktop login session** (the CLIs need the keychain). On macOS, a LaunchAgent:
|
|
158
|
+
|
|
159
|
+
```bash
|
|
160
|
+
cat > ~/Library/LaunchAgents/dev.omnilane.resign.plist <<'EOF'
|
|
161
|
+
<?xml version="1.0" encoding="UTF-8"?>
|
|
162
|
+
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
|
|
163
|
+
<plist version="1.0"><dict>
|
|
164
|
+
<key>Label</key><string>dev.omnilane.resign</string>
|
|
165
|
+
<key>ProgramArguments</key><array><string>/bin/zsh</string><string>-lc</string><string>omnilane resign</string></array>
|
|
166
|
+
<key>StartCalendarInterval</key><dict><key>Hour</key><integer>9</integer><key>Minute</key><integer>0</integer></dict>
|
|
167
|
+
<key>StandardOutPath</key><string>/tmp/omnilane-resign.log</string>
|
|
168
|
+
<key>StandardErrorPath</key><string>/tmp/omnilane-resign.log</string>
|
|
169
|
+
</dict></plist>
|
|
170
|
+
EOF
|
|
171
|
+
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/dev.omnilane.resign.plist
|
|
68
172
|
```
|
|
69
173
|
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
> itself. Its identity is read from the nearest launching CLI: explicit model and
|
|
75
|
-
> effort flags, or for Codex app-server/no-model launches, a bound current-turn
|
|
76
|
-
> rollout (never app-server startup defaults). An ordinary session passes nothing,
|
|
77
|
-
> and `omnilane whoami` prints that identity as a `--caller-context FILE`. With
|
|
78
|
-
> neither an assertion nor a readable identity, the dispatch is refused with
|
|
79
|
-
> `missing-caller-context` before any job is created.
|
|
174
|
+
This release verified `omnilane resign` from a desktop terminal session, including
|
|
175
|
+
a real unattended re-sign of a Codex self-update; the LaunchAgent wrapper above is
|
|
176
|
+
an example and was not itself exercised. Check it on your machine with
|
|
177
|
+
`launchctl kickstart gui/$(id -u)/dev.omnilane.resign`.
|
|
80
178
|
|
|
81
|
-
|
|
82
|
-
|
|
179
|
+
What `omnilane resign` exits with:
|
|
180
|
+
|
|
181
|
+
| Exit | Meaning | You do |
|
|
182
|
+
|---|---|---|
|
|
183
|
+
| 0 | nothing had changed, or everything that changed was re-signed | nothing |
|
|
184
|
+
| 10 | `--check` only: something changed | run `omnilane resign` |
|
|
185
|
+
| 20 | a vendor needs you: new or missing signer, unsigned or locally patched binary, new install directory, or the provider refused probes that passed last time | read the message. It prints either "retry later" or the exact `omnilane resign --vendor V --approve V` to run after you have looked |
|
|
186
|
+
| 30 | the re-signed overlay failed its real dispatch and the previous one was restored | nothing is broken; read the log |
|
|
187
|
+
| 2 | no overlay is configured | do Step 2 |
|
|
188
|
+
|
|
189
|
+
Two limits to know. Signer checks use macOS code signatures, so on Linux every
|
|
190
|
+
changed CLI stops at exit 20 for your `--approve`. And a binary with no real
|
|
191
|
+
signature (a locally patched CLI, for example) always stops for approval: nothing
|
|
192
|
+
ties it to its vendor, which is the point of the check.
|
|
193
|
+
|
|
194
|
+
### Step 4 — Try it from inside the assistant
|
|
195
|
+
|
|
196
|
+
Ask your assistant to run `omnilane whoami`. It should answer with the model and
|
|
197
|
+
effort it is running as, and a score. Then ask it to delegate something small:
|
|
198
|
+
"use omnilane to have the triage lane count the TODO comments in this repo".
|
|
199
|
+
|
|
200
|
+
If it is refused, the refusal says which check failed and what to do:
|
|
201
|
+
|
|
202
|
+
| `failed_gate` | In plain words | Fix |
|
|
203
|
+
|---|---|---|
|
|
204
|
+
| `caller-identity` | omnilane could not tell which model is asking | have it run `omnilane whoami` **as the only command** in that tool call. Codex in particular is unreadable behind `; echo $?`, `&&` or a pipe |
|
|
205
|
+
| `target-transport` | this machine has not proven that target, or a CLI updated since | `omnilane resign` (Step 3) |
|
|
206
|
+
| `downward-ceiling` | the target model scores higher than the model asking; a model may only delegate sideways or down | pick one of the `eligible_lanes` the refusal lists, or start the assistant at a higher effort |
|
|
207
|
+
|
|
208
|
+
A Codex automation that wakes an existing thread records no effort. omnilane
|
|
209
|
+
then holds that caller to its model's lowest score instead of refusing it: cheap
|
|
210
|
+
lanes keep working, expensive ones say which effort would reach them.
|
|
83
211
|
|
|
84
212
|
## 🧭 How it works
|
|
85
213
|
|
|
@@ -128,17 +256,17 @@ actually resolves.
|
|
|
128
256
|
|
|
129
257
|
| Lane | First choice | Backup | When |
|
|
130
258
|
|---|---|---|---|
|
|
131
|
-
| 🔥 hardest-coding | Claude Fable 5.1 (
|
|
132
|
-
| 🏗️ bulk-mechanical | GPT-5.6 Sol (high)
|
|
133
|
-
| 🧹 triage | GPT-5.6 Luna (high) | Gemini 3.8 Flash (Low) → Claude Haiku 4.5 | High-volume scans, first-pass filtering |
|
|
134
|
-
| ⚖️ hard-judgment | Claude Fable 5.1 (xhigh) | GPT-6 Astra (xhigh) → Grok 4.6 | Architecture arbitration, deep reasoning, second opinions |
|
|
135
|
-
| ✒️ taste-final | Claude Fable 5.1 (xhigh)
|
|
136
|
-
| 💬 consult | GPT-6 Astra (xhigh) | Claude Fable 5.1 (xhigh) → Grok 4.6 → Gemini 3.8 Flash (
|
|
137
|
-
| 🎨 ui-draft | GPT-
|
|
138
|
-
| 📚 long-context |
|
|
139
|
-
| ⚡ fast-agentic | Gemini 3.8 Flash (
|
|
140
|
-
| 📡 live-search | Grok 4.
|
|
141
|
-
| 🚰 coding-overflow | Grok 4.
|
|
259
|
+
| 🔥 hardest-coding | GPT-6 Astra (xhigh) | Claude Fable 5.1 (xhigh) → GPT-6 Astra (high) → Claude Fable 5.1 (high) → GPT-6 Astra (medium) → Claude Opus 5 (high) → GPT-6 Astra (low) → GPT-5.6 Sol (xhigh) → Grok 4.7 → Grok 4.6 → Gemini 3.8 Flash (High) | Hardest implementation, deep root-cause debug, correctness-critical edits |
|
|
260
|
+
| 🏗️ bulk-mechanical | GPT-6 Astra (low) | GPT-5.6 Sol (high) → Gemini 3.8 Flash (High) → Claude Opus 5 (medium) | Refactors, migrations, tests, review sweeps — mechanical endurance |
|
|
261
|
+
| 🧹 triage | GPT-5.6 Luna (high) | Gemini 3.8 Flash (Low) → Claude Sonnet 5 (low) → Claude Haiku 4.5 | High-volume scans, first-pass filtering |
|
|
262
|
+
| ⚖️ hard-judgment | Claude Fable 5.1 (xhigh) | GPT-6 Astra (xhigh) → Claude Opus 5 (max) → GPT-6 Astra (high) → Claude Opus 5 (xhigh) → Claude Opus 5 (high) → Grok 4.7 → Grok 4.6 → Gemini 3.8 Flash (High) | Architecture arbitration, deep reasoning, second opinions |
|
|
263
|
+
| ✒️ taste-final | Claude Opus 5 (max) | Claude Fable 5.1 (xhigh) → Claude Opus 5 (xhigh) → Grok 4.7 → GPT-6 Astra (xhigh) → Claude Opus 5 (high) → Grok 4.6 → Gemini 3.8 Flash (High) | User-facing prose and style arbitration; benchmarks do not prove visual or editorial taste |
|
|
264
|
+
| 💬 consult | GPT-6 Astra (xhigh) | Claude Fable 5.1 (xhigh) → Grok 4.7 → Grok 4.6 → Gemini 3.8 Flash (High) | Direct named-model consultation; keep `--vendor` to prevent fallback |
|
|
265
|
+
| 🎨 ui-draft | GPT-6 Astra (high) | Claude Opus 5 (high) → GPT-6 Astra (low) → Gemini 3.8 Flash (High) | UI drafts only with a design system or reference images; no aesthetic benchmark claim |
|
|
266
|
+
| 📚 long-context | Claude Opus 5 (high) | Claude Opus 5 (medium) → Claude Opus 5 (low) → GPT-5.6 Terra (max) → Gemini 3.8 Flash (High) | Long-document synthesis; context size alone does not prove task quality |
|
|
267
|
+
| ⚡ fast-agentic | GPT-6 Astra (low) | Gemini 3.8 Flash (Medium) → GPT-5.6 Sol (medium) → Claude Opus 5 (low) | Fast multi-step tool loops and multimodal checks |
|
|
268
|
+
| 📡 live-search | Grok 4.7 | Grok 4.6 → Gemini 3.8 Flash (High) → Claude Opus 5 (medium) | Realtime X/web search; backups provide generic web search, not equivalent X context |
|
|
269
|
+
| 🚰 coding-overflow | Grok 4.7 | Grok 4.6 → Gemini 3.8 Flash (High) → Kimi K3 → Qwen3 Coder Plus → OpenCode | Explicit Codex-quota relief; provider failure does not auto-retry another vendor |
|
|
142
270
|
| 🗳️ arbitrate | off (opt-in vote panel) | — | Built-in opinion panel for big calls — disabled by default; enable it in `routing.local.yaml`, one call per voter per round |
|
|
143
271
|
|
|
144
272
|
The **backup** is the next candidate in the lane's `routing.yaml` chain — what
|
|
@@ -160,52 +288,59 @@ request; this is not a free-form shell parser in `dispatch.sh`.
|
|
|
160
288
|
table. If an explicit target is absent or unavailable, the command fails
|
|
161
289
|
clearly instead of falling back to another vendor or family.
|
|
162
290
|
|
|
163
|
-
##
|
|
291
|
+
## Using the assistant's own sub-agents
|
|
164
292
|
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
On native rejection, auto reports a CLI reason and keeps the exact resolved
|
|
171
|
-
target rather than substituting another vendor/model.
|
|
293
|
+
By default omnilane hands work to a vendor's command-line tool. When the model
|
|
294
|
+
that should do the work belongs to the assistant's **own** vendor, going out
|
|
295
|
+
through a second CLI is a detour: another login, another process, another thing
|
|
296
|
+
that breaks when that CLI updates. Most assistants can start a sub-agent
|
|
297
|
+
themselves, and omnilane can plan the work for that instead. Two ways.
|
|
172
298
|
|
|
173
|
-
|
|
174
|
-
|
|
175
|
-
|
|
299
|
+
**A worker that runs exactly what the assistant runs (`--inherit`).** The
|
|
300
|
+
assistant starts its sub-agent *without choosing a model or an effort*, so the
|
|
301
|
+
worker is a copy of the caller. A copy cannot be stronger than the original,
|
|
302
|
+
which is the only thing omnilane's score check is there to prevent, so this path
|
|
303
|
+
needs no vendor CLI and no transport overlay, and it keeps working when the
|
|
304
|
+
caller's effort is unknown or an overlay has gone stale.
|
|
176
305
|
|
|
177
|
-
|
|
178
|
-
omnilane
|
|
179
|
-
|
|
306
|
+
```sh
|
|
307
|
+
omnilane native-context --workdir /absolute/repo --inherits-caller-runtime # prints a capability file
|
|
308
|
+
omnilane route --inherit --native-context /path/printed/above --workdir /absolute/repo triage "Count the TODO comments"
|
|
309
|
+
# -> a PENDING handoff (JSON). The assistant now starts its own sub-agent with no
|
|
310
|
+
# model argument, checks the result, and records it:
|
|
180
311
|
omnilane jobs --json complete-native JOB_ID /absolute/completion.json
|
|
181
312
|
omnilane jobs --json status JOB_ID
|
|
182
|
-
omnilane jobs --json result JOB_ID
|
|
183
|
-
omnilane jobs --json list --status pending
|
|
184
313
|
```
|
|
185
314
|
|
|
186
|
-
|
|
187
|
-
|
|
188
|
-
|
|
189
|
-
|
|
190
|
-
|
|
191
|
-
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
|
|
195
|
-
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
|
|
199
|
-
|
|
200
|
-
|
|
201
|
-
|
|
202
|
-
|
|
203
|
-
|
|
204
|
-
|
|
205
|
-
|
|
206
|
-
|
|
207
|
-
|
|
208
|
-
|
|
315
|
+
The honest part: the handoff is marked `satisfies_lane_target: false`. The lane
|
|
316
|
+
is only a label for the kind of work. A result produced this way was made by
|
|
317
|
+
"the assistant's own sub-agent", never by "the hardest-coding model", and a lane
|
|
318
|
+
that needs a stronger model than the caller is still refused.
|
|
319
|
+
`--inherits-caller-runtime` is the assistant's own statement that its sub-agent
|
|
320
|
+
tool behaves this way; omnilane cannot observe it. What is known per assistant:
|
|
321
|
+
|
|
322
|
+
| Assistant | Sub-agent without a model argument |
|
|
323
|
+
|---|---|
|
|
324
|
+
| Claude Code | documented to use the main conversation's model, and the session's effort unless the agent definition sets one. True for the built-in general-purpose agent with `CLAUDE_CODE_SUBAGENT_MODEL` unset. Run end to end in this release |
|
|
325
|
+
| Codex | `collaboration.spawn_agent` with no model and no effort. Run end to end in this release |
|
|
326
|
+
| Grok Build | documented to inherit the parent's model (the bundled `general-purpose` agent is `model: inherit`); effort not documented. Not run in this release |
|
|
327
|
+
| Antigravity | no sub-agent tool found in `agy` 1.2.7. Not available |
|
|
328
|
+
|
|
329
|
+
**A specific model the assistant's tool can select.** Describe what the tool
|
|
330
|
+
really accepts in a capability file (start from `omnilane native-context`, add a
|
|
331
|
+
row per exact model and effort) and pass `--native-context FILE` to an ordinary
|
|
332
|
+
`omnilane route`. omnilane uses the sub-agent only when a row matches exactly:
|
|
333
|
+
model, effort, mode, workdir, tools, isolation and lifecycle. Same vendor is not
|
|
334
|
+
same model, and nothing is guessed from the CLIs you have installed.
|
|
335
|
+
`--executor native` fails instead of falling back; `--executor cli` forces the
|
|
336
|
+
external CLI. When a same-vendor target goes out through the CLI only because no
|
|
337
|
+
file was given, dispatch now says so.
|
|
338
|
+
|
|
339
|
+
Either way the sub-agent shares the assistant's tools and filesystem: there is no
|
|
340
|
+
operating-system sandbox, and `advise`/`work` are intent, not enforcement.
|
|
341
|
+
Background, live, named-thread, multi-round, vote and `sysops` work stays on the
|
|
342
|
+
CLI path. Protocol handling needs Python 3.9+. Schemas, the completion file,
|
|
343
|
+
agent reuse and cancellation: [docs/native-executor.md](docs/native-executor.md).
|
|
209
344
|
|
|
210
345
|
<details>
|
|
211
346
|
<summary><b>Model-role guidance (delegation still required)</b></summary>
|
|
@@ -223,7 +358,7 @@ The commander orchestrates and validates; workers do not delegate again.
|
|
|
223
358
|
- **Codex · Sol** — delegate bulk-mechanical and constrained ui-draft at high. Escalate hardest coding and judgment to Fable/Astra; route long/fast work to Gemini 3.8 Flash and live search to Grok.
|
|
224
359
|
- **Codex · Astra** — prompt-level controller backup and independent reviewer. Use xhigh by default for hardest coding/judgment and consult/taste; explicitly select `--vendor codex --effort max` when needed. Explicit model/effort always win.
|
|
225
360
|
- **Codex · Terra** — delegate the Codex long-context fallback at max. Bulk stays on Sol high; escalate hard work to Fable/Astra.
|
|
226
|
-
- **Grok Build · Grok 4.6** — delegate live-search and coding-overflow, plus fallback duty in hardest-coding, hard-judgment, and taste-final. Dispatch primary hard coding/judgment/taste work to Codex/Claude/Gemini when available; verify API signatures and cited facts.
|
|
361
|
+
- **Grok Build · Grok 4.7 / 4.6** — delegate live-search and coding-overflow, plus fallback duty in hardest-coding, hard-judgment, and taste-final. Dispatch primary hard coding/judgment/taste work to Codex/Claude/Gemini when available; verify API signatures and cited facts.
|
|
227
362
|
- **Antigravity · Gemini 3.8 Flash** — delegate long-context Medium, fast-agentic/triage Low, and bulk/overflow/web fallbacks High. Do not infer visual taste or controller authority from agent/coding benchmarks.
|
|
228
363
|
|
|
229
364
|
</details>
|
|
@@ -637,6 +772,9 @@ When that fails, run `omnilane whoami` — it either prints a
|
|
|
637
772
|
`--effort`, a model alias, no scored row). A model must not assert the human
|
|
638
773
|
exemption for itself.
|
|
639
774
|
|
|
775
|
+
Every refusal is one JSON line on stderr. Read `failed_gate`, `reason` and
|
|
776
|
+
`next_command` first; `eligible_lanes` lists what you can still dispatch.
|
|
777
|
+
|
|
640
778
|
`runtime-mapping-unverified` — your identity is fine, but the *target* has no
|
|
641
779
|
proven host-local request selector. Either it was never probed, or its probe
|
|
642
780
|
failed; `omnilane doctor` reports the count of such configurations and the
|
|
@@ -650,7 +788,13 @@ and runner-script hash, and Codex and Claude evidence paths embed version
|
|
|
650
788
|
directories, so an upgrade removes the file rather than changing its digest.
|
|
651
789
|
Tagged evidence degrades only its own vendor; untagged evidence, such as the
|
|
652
790
|
probe manifest, still closes the whole gate. Doctor names the file and the
|
|
653
|
-
vendor;
|
|
791
|
+
vendor; `omnilane resign` re-probes and re-signs it, and the dispatch skill
|
|
792
|
+
carries the runbook behind that command.
|
|
793
|
+
|
|
794
|
+
`target-above-effective-ceiling` — nothing is broken. The target scores above the
|
|
795
|
+
caller. `required_caller_effort` names the effort the calling session would need;
|
|
796
|
+
a caller marked `caller_degraded` was launched by a harness that recorded no
|
|
797
|
+
effort (a Codex heartbeat automation does this) and is held to its model's floor.
|
|
654
798
|
|
|
655
799
|
</details>
|
|
656
800
|
|
|
@@ -706,6 +850,108 @@ working notes, including per-benchmark caveats, live in
|
|
|
706
850
|
|
|
707
851
|
## 📜 Release history
|
|
708
852
|
|
|
853
|
+
## What's new in v0.45.0
|
|
854
|
+
|
|
855
|
+
- **Run `omnilane resign` once after upgrading.** The score registry moved to a
|
|
856
|
+
new snapshot and the transport overlay is bound to it; until the overlay is
|
|
857
|
+
rebuilt, a model caller is refused on every lane with
|
|
858
|
+
`transport overlay snapshot mismatch`.
|
|
859
|
+
- **Scores follow Artificial Analysis Intelligence Index v4.3.2.** v4.2 and
|
|
860
|
+
v4.3.2 are different scales with uneven gaps (Fable 5.1 max 57 → 53, Grok 4.6
|
|
861
|
+
high 51 → 44, Sol high 48 → 42), so every row was re-scored rather than one
|
|
862
|
+
added. Ceilings move with it: Fable 5.1 max, Fable 5.1 xhigh and Astra max now
|
|
863
|
+
tie at 53, and a mid-effort controller that reached nothing in `hard-judgment`
|
|
864
|
+
before now does. `scripts/aa_rebaseline.py` rebuilds the registry from a saved
|
|
865
|
+
AA extract, so the next index revision is a re-run.
|
|
866
|
+
- **The lane table is rewritten, not patched.** Each lane now names the
|
|
867
|
+
measurements that match its kind of work and lists candidates best-first on
|
|
868
|
+
them: hard terminal coding (Terminal-Bench 4.0, the coding benchmark the
|
|
869
|
+
frontier has not saturated) for `hardest-coding`, expert-level reasoning and
|
|
870
|
+
graded analytical quality for `hard-judgment`, graded document presentation for
|
|
871
|
+
`taste-final`, automation accuracy per minute for `fast-agentic`, the harder
|
|
872
|
+
long-context benchmark for `long-context`, visual understanding for `ui-draft`.
|
|
873
|
+
Eight lanes change their first choice: Astra (xhigh) leads `hardest-coding`,
|
|
874
|
+
Astra (low) leads `bulk-mechanical` and `fast-agentic`, Astra (high) leads
|
|
875
|
+
`ui-draft`, Opus 5 (max) leads `taste-final`, Opus 5 (high) leads
|
|
876
|
+
`long-context`, and Grok 4.7 leads `coding-overflow` as well as `live-search`.
|
|
877
|
+
Every chain also steps down through the score range, so whatever a
|
|
878
|
+
controller's ceiling, the first candidate it can reach is the best one it can
|
|
879
|
+
reach. Fable max and Astra max are in no chain: xhigh is level or ahead for
|
|
880
|
+
markedly less. The figures are in `docs/model-capabilities-2026-09.md`.
|
|
881
|
+
- **Grok 4.7 joins every lane Grok serves**, ahead of Grok 4.6, which stays
|
|
882
|
+
behind it: a lane skips a candidate this host has not proven, so 4.6 keeps
|
|
883
|
+
serving until `resign` has probed 4.7. With `--vendor grok` there is no
|
|
884
|
+
fall-through; until then ask for `--model grok-4.6`.
|
|
885
|
+
- **A dead fallback fixed.** `claude claude-sonnet-5 high` could never be
|
|
886
|
+
dispatched under the old registry; it now resolves once probed.
|
|
887
|
+
- Upgrade: `npm i -g omnilane@0.45.0`, then `omnilane resign`. For Grok 4.7 the
|
|
888
|
+
`grok` CLI must be logged in when you re-sign.
|
|
889
|
+
|
|
890
|
+
## What's new in v0.44.0
|
|
891
|
+
|
|
892
|
+
- **A CLI you patch yourself is re-signed unattended too.** If a local step
|
|
893
|
+
re-signs a vendor CLI adhoc after every update (a post-update patch, for
|
|
894
|
+
instance), the signer check used to hold every such update for `--approve`.
|
|
895
|
+
Run `omnilane resign --trust-adhoc VENDOR` once per vendor and an adhoc update
|
|
896
|
+
in the same install directory now goes through the daily `omnilane resign`
|
|
897
|
+
like a same-signer one. An unsigned executable, an adhoc one in another
|
|
898
|
+
directory, and every other vendor still stop for you. The trust is recorded
|
|
899
|
+
on the overlay, survives later re-signs of any vendor, and is an operator
|
|
900
|
+
action a model never runs. Works for all four vendors.
|
|
901
|
+
- **An expired login says "log in", not "retry later".** `Failed to
|
|
902
|
+
authenticate`, `OAuth session expired`, `Invalid API key`, `Unauthorized` and
|
|
903
|
+
`401` now mark the vendor unprobeable, and the held-vendor message tells you
|
|
904
|
+
to log in first.
|
|
905
|
+
- Upgrade: `npm i -g omnilane@0.44.0`. If you are coming from 0.42.x, run
|
|
906
|
+
`omnilane resign --record-signers` once as well (see the 0.43.0 notes).
|
|
907
|
+
|
|
908
|
+
## What's new in v0.43.1
|
|
909
|
+
|
|
910
|
+
Install this rather than 0.43.0. In 0.43.0, `build_overlay.py` and `probe.py`
|
|
911
|
+
failed to import on Python 3.9, which broke the first-install overlay build and
|
|
912
|
+
`omnilane resign` on that version. Nothing else changed; everything in the 0.43.0
|
|
913
|
+
notes below applies. Upgrade: `npm i -g omnilane@0.43.1`, then once:
|
|
914
|
+
`omnilane resign --record-signers`.
|
|
915
|
+
|
|
916
|
+
## What's new in v0.43.0
|
|
917
|
+
|
|
918
|
+
In ten days 0.42.x refused every model caller four times, each time over a fact
|
|
919
|
+
omnilane does not control: a renamed launcher, a runner script changed without a
|
|
920
|
+
re-sign, four vendor CLIs updating themselves in one week, and a Codex automation
|
|
921
|
+
that records no effort. Each became "nothing can be dispatched". This release
|
|
922
|
+
turns each into a narrower, explained outcome, and repairs the common one itself.
|
|
923
|
+
|
|
924
|
+
- **Vendor CLI updated? `omnilane resign`.** It finds what changed, re-probes only
|
|
925
|
+
that vendor, checks the result, swaps it in, confirms with one real dispatch,
|
|
926
|
+
and restores the old overlay if that fails. It re-signs **unattended** only when
|
|
927
|
+
the new executable carries the same code-signing team and sits in the same
|
|
928
|
+
place; anything else stops with the exact `--approve` command for you. Run
|
|
929
|
+
`omnilane resign --record-signers` once, schedule `omnilane resign` daily, and
|
|
930
|
+
updates stop being your problem. Verified on a real Codex self-update
|
|
931
|
+
(0.155.0 → 0.155.1): no approval, every mapping kept, exit 0.
|
|
932
|
+
- **A refusal tells the model what to do.** Every refused dispatch now carries
|
|
933
|
+
`failed_gate`, `reason`, `next_command`, `required_caller_effort`, and
|
|
934
|
+
`eligible_lanes` — the lanes that caller *can* reach right now.
|
|
935
|
+
- **No recorded effort narrows instead of refusing.** A Codex heartbeat
|
|
936
|
+
automation is held to its model's lowest score rather than being refused on
|
|
937
|
+
every lane. Cheap lanes keep working; expensive ones say which effort is needed.
|
|
938
|
+
- **The assistant's own sub-agents.** `omnilane native-context` writes the
|
|
939
|
+
capability file that used to be hand-made, and `omnilane route --inherit` plans
|
|
940
|
+
a worker that is a copy of the caller: no external CLI, no overlay, works even
|
|
941
|
+
when the caller cannot be identified, and is honestly marked as *not* the
|
|
942
|
+
lane's target model. Run end to end in Claude Code and Codex desktop.
|
|
943
|
+
- **Codex: one omnilane command per tool call.** `omnilane whoami; echo $?`
|
|
944
|
+
cannot be identified, `omnilane whoami` alone can. The refusal now says so.
|
|
945
|
+
- **Doctor sees a CLI that moved** beside its old file, and warns, with the
|
|
946
|
+
steps, when no overlay exists at all.
|
|
947
|
+
- **Rewritten skill and tutorial.** The skill is now a five-step procedure a
|
|
948
|
+
model follows; this README walks through letting an assistant drive omnilane.
|
|
949
|
+
- **Limits.** Unattended re-signing relies on macOS code signatures; on Linux, and
|
|
950
|
+
for any unsigned or locally patched CLI, every update asks for `--approve`.
|
|
951
|
+
`--inherit` has not been run inside Grok Build, and Antigravity exposes no
|
|
952
|
+
sub-agent tool. Full detail: [CHANGELOG](CHANGELOG.md).
|
|
953
|
+
- **Upgrade.** `npm i -g omnilane@0.43.0`, then once: `omnilane resign --record-signers`.
|
|
954
|
+
|
|
709
955
|
## What's new in v0.42.9
|
|
710
956
|
|
|
711
957
|
- **Codex desktop behind a launcher.** When ChatGPT.app starts its app-server
|