model-orchestrator 0.1.12 → 0.1.14
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +35 -1
- package/README.md +58 -2
- package/bin/cli-run.mjs +144 -19
- package/bin/cli.js +9 -4
- package/docs/part-1-beginner.md +2 -0
- package/docs/part-2-intermediate.md +7 -1
- package/package.json +7 -3
- package/src/install.js +31 -1
- package/templates/advanced/vm/README.md +2 -1
- package/templates/agents/agy/README.md +1 -1
- package/templates/agents/agy/finding-verifier.md +29 -0
- package/templates/agents/claude-code/README.md +2 -1
- package/templates/agents/claude-code/finding-verifier.md +43 -0
- package/templates/agents/snippets/claude-code.md +1 -0
- package/templates/common/README.md +1 -4
- package/templates/common/protocols/build-protocol.md +11 -1
- package/templates/intermediate/CLI-RUN.md +40 -3
- package/templates/intermediate/ROUTING.md +6 -1
- package/templates/intermediate/TIERS.md +42 -1
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,38 @@ All notable changes to this project are documented here. The format follows [Kee
|
|
|
4
4
|
|
|
5
5
|
## [Unreleased]
|
|
6
6
|
|
|
7
|
+
## [0.1.14] - 2026-09-09
|
|
8
|
+
|
|
9
|
+
Three refinements to the routing model, from a review by [@shawnwows](https://x.com/shawnwows). The theme is the same in all three: a routing decision that was implied, inherited or asserted is now stated, pinned or checked.
|
|
10
|
+
|
|
11
|
+
### Added
|
|
12
|
+
|
|
13
|
+
- **`--model` and `--effort` on every lane, and a route recorded per run.** A lane with no flag and no `defaults` entry in `bin/lanes.json` runs on its own config file, which `cli-run` cannot see: a CLI configured months ago at a low reasoning effort keeps auditing at that effort while the routing docs describe an adversarial pass, and nothing raises an error. Each vendor spells the flags differently and `cli-run` translates (`grok -m/--reasoning-effort`, `codex -m/-c model_reasoning_effort="X"`, `agy --model/--effort`, `hermes -m/--reasoning`, `qwen -m` and no reasoning flag), each one read from that CLI's own `--help`. Flags beat `defaults`, `defaults` beats nothing, `--doctor` prints what each lane is pinned to, and the log carries `model_requested`, `effort_requested`, `model_source` and `effort_source` on every record, including runs refused before the lane started. It records no "actual": one lane of five (grok) reports a model id in its own output and the other four report none, so the field would be populated for one lane and empty for four, and it would be a provider-supplied string, which the durable log never holds. `--effort` on qwen is a usage error rather than a silent drop, and route values are charset-bounded because a model id becomes an argv element and, on codex, part of a TOML value.
|
|
14
|
+
- **`finding-verifier`, a sixth subagent, in both agent formats.** Review and scanner findings no longer go straight to a repair. It reads the cited line, states what would trigger the problem, hunts for the guard, caller or test that makes it impossible, and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Only CONFIRMED earns a change; INCONCLUSIVE is never rounded up to be safe or down to be tidy. Bound into the build protocol as Stage 5a, into `ROUTING.md`, and into the Claude Code activation snippet. The reproduction rule already existed in Stage 5; it had no owner, no separate model family and no way to say "I could not settle this".
|
|
15
|
+
- **Complexity and risk as inputs, alongside role** (`TIERS.md`). Complexity moves the effort: a worker executing a finished plan needs less reasoning than the reviewer judging its output. Risk (security, privacy, data loss, irreversible) moves the tier and who reads the result, because none of those failures is fixable by editing the code afterwards. A one-line change to an auth check is simple and high-risk at once, and the risk decides. Deliberately two rules and two small tables rather than a role by complexity by risk matrix: an 80-cell table is not maintained, and an unmaintained routing table is worse than none because it is believed.
|
|
16
|
+
|
|
17
|
+
### Changed
|
|
18
|
+
|
|
19
|
+
- `--model` is no longer qwen-only. `--safe-mode` still is.
|
|
20
|
+
- The route is resolved before the "lane disabled" and "binary missing" refusals, so those records carry it too. Found by the pre-release audit: a run refused for a missing binary is still a run that requested a route, and a failure record without one is the gap this release exists to close.
|
|
21
|
+
- `bin/lanes.json` gains an optional `defaults` block. It fails closed with the rest of the file: an unknown lane, an unknown key, a value outside the charset, or an effort pinned on a lane with no reasoning flag refuses every lane until it is fixed, rather than being skipped quietly.
|
|
22
|
+
- The generated activation list gains a step about pinning the route, and `--doctor` output gains a route column with a plain sentence about what "not pinned" means.
|
|
23
|
+
|
|
24
|
+
## [0.1.13] - 2026-09-08
|
|
25
|
+
|
|
26
|
+
Three issues from a fresh first-run walkthrough of 0.1.12 (#26, #27, #28). Same class as 0.1.12's five: a surface describing an install that did not happen. A fourth, #25, was filed and closed as a mistake on the reporter's side, not a defect: the warning it said was missing has been printed since 0.1.12 and the repro had been read through a truncated pipe.
|
|
27
|
+
|
|
28
|
+
### Fixed
|
|
29
|
+
|
|
30
|
+
- **The "Then prove it took" list no longer sends a level 1 reader to a file level 1 never wrote** (#27). Step 4 told every reader, at every level, to pick a lane out of `bin/lanes.json` and run `node bin/cli-run.mjs`. Level 1 writes no `bin/` at all, and step 3 immediately above it hedged correctly with "At level 2+" while step 4 did not. The list is now `proofSteps()` in `src/install.js`, gated on level the same way `activationSteps()` is, and the template renders it. Two tests: the README's section must equal the array exactly for every level and primary, and no `bin/` path may appear in it that the plan did not write.
|
|
31
|
+
- **The box setup no longer tells you to sign in to CLIs you did not pick** (#26). `templates/advanced/vm/README.md` step 3 was a fixed sentence naming `codex login --device-auth`, `grok login --device-auth` and `agy`. A level 3 install of claude-code, codex, qwen and ollama was told to sign in to two CLIs it does not have and never told about the one it does. The step now renders each selected CLI's own `auth` string from the catalog. Everything else in that file was already computed from the selection, which is what made the one hardcoded line easy to miss.
|
|
32
|
+
- **A selected local runtime is finally told to install itself** (#26). `activationSteps()` filtered on `kind === 'agent-cli'`, so Ollama, which has a binary and a download page, appeared in no ordered list at any level. Its only mention was one row of a URL table in `DELEGATION_MATRIX.md`. It now gets a step naming the download page and the `ollama pull <model>` that has to follow it.
|
|
33
|
+
- **The tool block stopped saying the same word twice** (#28). Every run that selected a tool printed `optional: Optional. Needs Python 3.10+ and uv.`, because the label repeated the note's own first word. The label is `note:` now. The note keeps the word, because `--list` and the interactive picker print it bare with no label.
|
|
34
|
+
|
|
35
|
+
### Changed
|
|
36
|
+
|
|
37
|
+
- `--primary` is documented as what it is. `--help` called it "required when several qualify", and then a `--yes` run with several candidates silently picked one in catalog order. The run now names the choice in the plan (`primary claude-code (chosen for you from claude-code, codex; pass --primary to decide it yourself)`) and the help says the same thing. Behaviour is unchanged: the default was sensible, only the promise was wrong.
|
|
38
|
+
|
|
7
39
|
## [0.1.12] - 2026-09-08
|
|
8
40
|
|
|
9
41
|
Five issues from one first-run walkthrough of 0.1.11 (#20 to #24). Every one of them is the same failure: a page describing an install that did not happen. Each fix removes the second copy of a fact rather than correcting it.
|
|
@@ -185,7 +217,9 @@ First release.
|
|
|
185
217
|
- Tests: a case per fix, judges proven to go red, mutation checks; `npm test` prints the current count.
|
|
186
218
|
- Adversarial audit: two Codex rounds plus a two-engine review (Codex, Antigravity); findings and fixes in `docs/audit-brief.md`. After the review: subagents go to the project root (`--project`), snippet paths computed from `--dir`, lane sections rendered from the selection, a primary agent required, level 3 asks for API keys separately from CLIs, images and CLI installs pinned, an activation summary at the end of every install.
|
|
187
219
|
|
|
188
|
-
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.
|
|
220
|
+
[Unreleased]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.14...HEAD
|
|
221
|
+
[0.1.14]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.13...v0.1.14
|
|
222
|
+
[0.1.13]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.12...v0.1.13
|
|
189
223
|
[0.1.12]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.11...v0.1.12
|
|
190
224
|
[0.1.11]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.10...v0.1.11
|
|
191
225
|
[0.1.10]: https://github.com/aunysillyme/model-orchestrator/compare/v0.1.9...v0.1.10
|
package/README.md
CHANGED
|
@@ -25,7 +25,7 @@ It never writes a secret, never runs a vendor shell script for you, and never ov
|
|
|
25
25
|
| Level | You have | You get |
|
|
26
26
|
|---|---|---|
|
|
27
27
|
| **1 · Beginner** | one LLM or one agent | tiers, task classification, the two build checkpoints, the protocols (build, propagate, gap analysis, deep research, numbers and logic, memory and record), a task-bundle template, and your agent set up to follow them |
|
|
28
|
-
| **2 · Intermediate** | several AIs with CLIs | everything above, plus `cli-run` (exit 0 means a structurally accepted non-empty response; opt-in `--expect-file` / `--expect-json` for real contracts), a delegation matrix generated from your selection, research triage across the lanes you have |
|
|
28
|
+
| **2 · Intermediate** | several AIs with CLIs | everything above, plus `cli-run` (exit 0 means a structurally accepted non-empty response; opt-in `--expect-file` / `--expect-json` for real contracts; `--model` / `--effort` to pin the route and log it), a delegation matrix generated from your selection, research triage across the lanes you have |
|
|
29
29
|
| **3 · Advanced** | a virtual machine | everything above, plus a gateway config rendered from the API keys you hold (asked separately from your CLIs), pinned images, box rules, privacy gates, and a weekly gap-analysis job with "what watches it" written down |
|
|
30
30
|
|
|
31
31
|
Read the thinking behind each level in [docs/](docs/README.md): [Part 1](docs/part-1-beginner.md) · [Part 2](docs/part-2-intermediate.md) · [Part 3](docs/part-3-advanced.md).
|
|
@@ -90,7 +90,7 @@ ai-orchestrator/
|
|
|
90
90
|
TASK_BUNDLE.md the brief every delegation carries
|
|
91
91
|
protocols/ build-protocol · propagate · gap-analysis · deep-research · numbers-and-logic · memory-and-record
|
|
92
92
|
CODECALC.md OBSIDIAN-TC.md mcp/ companion-tool install docs + per-agent registration snippets (if selected)
|
|
93
|
-
<project>/.claude/agents/
|
|
93
|
+
<project>/.claude/agents/ six subagents, one per tier plus finding-verifier, at the PROJECT root (if Claude Code is primary)
|
|
94
94
|
CLAUDE.snippet.md the block to paste into your CLAUDE.md
|
|
95
95
|
ROUTING.md multi-lane decision tree (level 2+)
|
|
96
96
|
TIERS.md DELEGATION_MATRIX.md RESEARCH_TRIAGE.md CLI-RUN.md
|
|
@@ -162,6 +162,54 @@ It deliberately does not run in this repository's CI. A canary is only meaningfu
|
|
|
162
162
|
6. **The orchestrator owns the main build.** Delegates hold none of your rules; they get bounded sub-parts and a brief.
|
|
163
163
|
7. **Only one process holds keys.** Names in the environment, values in a secrets manager, never in a file here.
|
|
164
164
|
|
|
165
|
+
## Routing by role, complexity and risk
|
|
166
|
+
|
|
167
|
+
Role picks the agent. Two more inputs move the choice, and they move it in
|
|
168
|
+
different directions, so `TIERS.md` states them separately rather than folding
|
|
169
|
+
them into the role:
|
|
170
|
+
|
|
171
|
+
- **Complexity moves the effort.** A worker executing a finished plan needs less
|
|
172
|
+
reasoning than the reviewer judging its output. When the plan is airtight the
|
|
173
|
+
spec is carrying the thinking.
|
|
174
|
+
- **Risk moves the tier and the reader.** Security, privacy, data loss and
|
|
175
|
+
irreversible changes buy the attack lane, a named check, a rollback path or a
|
|
176
|
+
human yes. A one-line change to an auth check is simple and high-risk at the
|
|
177
|
+
same time, and it is the risk that decides.
|
|
178
|
+
|
|
179
|
+
The top of the ladder is bought with evidence: a reproduced failure, an
|
|
180
|
+
unresolved checkpoint, an irreversible change. A task that merely feels hard is
|
|
181
|
+
a deep-tier task, not an escalation.
|
|
182
|
+
|
|
183
|
+
## A finding is a claim, not a fact
|
|
184
|
+
|
|
185
|
+
Review findings do not go straight to a repair. `finding-verifier` reads the
|
|
186
|
+
cited line, states what would trigger the problem, then hunts for the guard,
|
|
187
|
+
caller or test that makes it impossible, and returns **CONFIRMED**,
|
|
188
|
+
**NOT_REPRODUCED** or **INCONCLUSIVE** per finding. Only CONFIRMED earns a
|
|
189
|
+
change. Use a different model family from the one that produced the finding
|
|
190
|
+
where you have one: a family asked to check its own claim tends to agree with
|
|
191
|
+
itself.
|
|
192
|
+
|
|
193
|
+
## Pin the route, or know that you did not
|
|
194
|
+
|
|
195
|
+
A lane with no `--model`, no `--effort` and no `defaults` entry in
|
|
196
|
+
`bin/lanes.json` runs on **its own config file**, which `cli-run` cannot see. A
|
|
197
|
+
CLI configured months ago at a low reasoning effort keeps auditing at that
|
|
198
|
+
effort while your routing docs describe an adversarial pass.
|
|
199
|
+
|
|
200
|
+
```bash
|
|
201
|
+
node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
|
|
202
|
+
node bin/cli-run.mjs --doctor # prints what each lane is pinned to, and what is not pinned
|
|
203
|
+
```
|
|
204
|
+
|
|
205
|
+
Every run logs the model and effort **requested** and where the request came
|
|
206
|
+
from: `flag`, `lanes.json`, or `lane_default`, on every record including the
|
|
207
|
+
runs that never reached a lane. It does not log an actual. One lane of five
|
|
208
|
+
(grok) reports a model id in its own output and the other four report none, so
|
|
209
|
+
an actual field would be present for one lane and missing for four, and it
|
|
210
|
+
would be a provider-supplied string, which the durable log deliberately never
|
|
211
|
+
holds.
|
|
212
|
+
|
|
165
213
|
## Requirements
|
|
166
214
|
|
|
167
215
|
Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box templates assume Ubuntu. Windows is untested: `cli-run` ends a lane's process tree there with `taskkill`, but nothing in CI runs on Windows, so treat it as unsupported until someone reports otherwise.
|
|
@@ -177,6 +225,14 @@ Node 18 or newer. No dependencies. Works on macOS and Linux; the level 3 box tem
|
|
|
177
225
|
|
|
178
226
|
Add an AI to `src/catalog.js` and every prompt, table, config and doc picks it up. Run `npm test`. Keep templates free of logic and free of anything that looks like a credential. The rest is in [CONTRIBUTING.md](CONTRIBUTING.md); releases in [RELEASING.md](RELEASING.md); security reports in [SECURITY.md](SECURITY.md).
|
|
179
227
|
|
|
228
|
+
## Credits
|
|
229
|
+
|
|
230
|
+
- [@shawnwows](https://x.com/shawnwows) reviewed the router and made the case for
|
|
231
|
+
separating role, complexity and risk instead of compressing them into one
|
|
232
|
+
scale, for recording the model and effort a lane was actually asked for, and
|
|
233
|
+
for verifying findings before they trigger repairs. All three shipped in
|
|
234
|
+
0.1.14.
|
|
235
|
+
|
|
180
236
|
## License
|
|
181
237
|
|
|
182
238
|
[MIT](LICENSE)
|
package/bin/cli-run.mjs
CHANGED
|
@@ -39,6 +39,17 @@
|
|
|
39
39
|
//
|
|
40
40
|
// The durable log stores a FIXED reason code per run (see REASONS), never a
|
|
41
41
|
// provider-supplied string. Bounded vendor stderr goes to your terminal only.
|
|
42
|
+
//
|
|
43
|
+
// ROUTE: which model and reasoning effort a lane ran with.
|
|
44
|
+
// A lane with no --model and no lanes.json default inherits whatever its own
|
|
45
|
+
// config file says, which is invisible from here and is how a documented route
|
|
46
|
+
// silently stops being the route that runs. --model / --effort pin it per call,
|
|
47
|
+
// `defaults` in lanes.json pins it per lane, and every run logs the value that
|
|
48
|
+
// was REQUESTED plus where the request came from (flag, lanes.json, or nothing
|
|
49
|
+
// at all). It does not log an "actual". One lane of five (grok) does report a
|
|
50
|
+
// model id in its own output; the other four report none, and a field present
|
|
51
|
+
// for one lane and absent for four is worse than no field. It would also be a
|
|
52
|
+
// provider-supplied string, which this log deliberately never holds.
|
|
42
53
|
|
|
43
54
|
import { spawn } from 'node:child_process';
|
|
44
55
|
import { StringDecoder } from 'node:string_decoder';
|
|
@@ -190,28 +201,71 @@ export function judgeQwen(rc, out) {
|
|
|
190
201
|
return pass(text, `subtype=success, totalErrors=0 across ${Object.keys(models).length} model(s)`);
|
|
191
202
|
}
|
|
192
203
|
|
|
204
|
+
// --- route: model and effort per lane --------------------------------------
|
|
205
|
+
// Each vendor spells these differently, and the spelling was read from each
|
|
206
|
+
// CLI's own --help, not remembered. A lane with `effort: null` has no reasoning
|
|
207
|
+
// flag at all; asking for one there is a usage error, never a silent drop.
|
|
208
|
+
// grok -m MODEL --reasoning-effort EFFORT
|
|
209
|
+
// codex -m MODEL -c model_reasoning_effort="EFFORT" (a TOML override, hence the quotes)
|
|
210
|
+
// agy --model M --effort EFFORT (low|medium|high)
|
|
211
|
+
// hermes -m MODEL --reasoning LEVEL (none|minimal|...)
|
|
212
|
+
// qwen -m MODEL no reasoning flag
|
|
213
|
+
export const LANE_FLAGS = {
|
|
214
|
+
grok: { model: (v) => ['-m', v], effort: (v) => ['--reasoning-effort', v] },
|
|
215
|
+
codex: { model: (v) => ['-m', v], effort: (v) => ['-c', `model_reasoning_effort="${v}"`] },
|
|
216
|
+
agy: { model: (v) => ['--model', v], effort: (v) => ['--effort', v] },
|
|
217
|
+
hermes: { model: (v) => ['-m', v], effort: (v) => ['--reasoning', v] },
|
|
218
|
+
qwen: { model: (v) => ['-m', v], effort: null }
|
|
219
|
+
};
|
|
220
|
+
|
|
221
|
+
// A model id or effort level becomes an argv element and, for codex, part of a
|
|
222
|
+
// TOML value. Bounding the charset is what makes both safe: no leading dash (a
|
|
223
|
+
// value cannot become a flag), no quote, space or control character (a value
|
|
224
|
+
// cannot break out of the TOML string), and a length cap so a config file
|
|
225
|
+
// cannot push an unbounded string into the durable log.
|
|
226
|
+
export const ROUTE_VALUE = /^[A-Za-z0-9][A-Za-z0-9._:@/+-]{0,63}$/;
|
|
227
|
+
export function badRouteValue(kind, v) {
|
|
228
|
+
if (typeof v !== 'string' || !ROUTE_VALUE.test(v)) {
|
|
229
|
+
return `--${kind} must be 1 to 64 characters of letters, digits, dot, underscore, colon, at, slash, plus or dash, and may not start with a dash: ${JSON.stringify(v)}`;
|
|
230
|
+
}
|
|
231
|
+
return null;
|
|
232
|
+
}
|
|
233
|
+
|
|
193
234
|
// --- adapters: build argv for a lane -------------------------------------
|
|
235
|
+
// Route flags go in front of the prompt for every lane, because two lanes
|
|
236
|
+
// (hermes, codex) take the prompt as a positional argument and a flag after it
|
|
237
|
+
// is either ignored or read as part of it.
|
|
238
|
+
function routeFlags(lane, opts) {
|
|
239
|
+
const spec = LANE_FLAGS[lane];
|
|
240
|
+
const out = [];
|
|
241
|
+
if (!spec) return out;
|
|
242
|
+
if (opts.model) out.push(...spec.model(opts.model));
|
|
243
|
+
if (opts.effort && spec.effort) out.push(...spec.effort(opts.effort));
|
|
244
|
+
return out;
|
|
245
|
+
}
|
|
246
|
+
|
|
194
247
|
export function buildArgv(lane, binary, prompt, opts, tmp) {
|
|
195
248
|
const timeout = opts.timeout;
|
|
249
|
+
const route = routeFlags(lane, opts);
|
|
196
250
|
switch (lane) {
|
|
197
251
|
case 'grok':
|
|
198
|
-
return { argv: [binary, '--output-format', 'json', '-p', prompt] };
|
|
252
|
+
return { argv: [binary, '--output-format', 'json', ...route, '-p', prompt] };
|
|
199
253
|
case 'codex': {
|
|
200
254
|
const last = join(tmp, 'last.txt');
|
|
201
255
|
const argv = [binary, 'exec', '--json', '--color', 'never', '--skip-git-repo-check', '-o', last];
|
|
202
256
|
if (opts.audit) argv.push('--sandbox', 'read-only'); // an audit lane that can write is a bug
|
|
257
|
+
argv.push(...route);
|
|
203
258
|
argv.push(prompt);
|
|
204
259
|
return { argv, outFile: last };
|
|
205
260
|
}
|
|
206
261
|
case 'agy': {
|
|
207
262
|
const mins = Math.max(1, Math.round(timeout / 60));
|
|
208
|
-
return { argv: [binary, '--print-timeout', `${mins}m`, '--output-format', 'stream-json', '-p', prompt] };
|
|
263
|
+
return { argv: [binary, '--print-timeout', `${mins}m`, '--output-format', 'stream-json', ...route, '-p', prompt] };
|
|
209
264
|
}
|
|
210
265
|
case 'hermes':
|
|
211
|
-
return { argv: [binary, '-z', prompt, '--usage-file', join(tmp, 'usage.json')] };
|
|
266
|
+
return { argv: [binary, '-z', ...route, prompt, '--usage-file', join(tmp, 'usage.json')] };
|
|
212
267
|
case 'qwen': {
|
|
213
|
-
const argv = [binary, '-o', 'json'];
|
|
214
|
-
if (opts.model) argv.push('-m', opts.model);
|
|
268
|
+
const argv = [binary, '-o', 'json', ...route];
|
|
215
269
|
if (opts.safeMode) argv.push('--safe-mode');
|
|
216
270
|
argv.push('-p', prompt);
|
|
217
271
|
return { argv };
|
|
@@ -370,26 +424,67 @@ function log(rec) {
|
|
|
370
424
|
// documented default). PRESENT BUT UNREADABLE OR MALFORMED = no lane enabled:
|
|
371
425
|
// a half-written config must fail closed, never re-enable what the installer
|
|
372
426
|
// disabled. Returns null when the file is bad so the caller can say so.
|
|
373
|
-
export function
|
|
427
|
+
export function laneConfig(here = dirname(fileURLToPath(import.meta.url))) {
|
|
374
428
|
const p = join(here, 'lanes.json');
|
|
375
|
-
if (!existsSync(p)) return LANES;
|
|
429
|
+
if (!existsSync(p)) return { enabled: LANES, defaults: {} };
|
|
376
430
|
try {
|
|
377
431
|
const j = JSON.parse(readFileSync(p, 'utf8'));
|
|
378
432
|
if (!j || typeof j !== 'object' || !Array.isArray(j.enabled)) return null;
|
|
379
433
|
if (!j.enabled.every((l) => typeof l === 'string' && LANES.includes(l))) return null;
|
|
380
|
-
|
|
434
|
+
// `defaults` pins a model and effort per lane. It is optional; present and
|
|
435
|
+
// malformed fails closed with the rest of the file, because a half-written
|
|
436
|
+
// route is exactly the silent-inheritance problem this field exists to fix.
|
|
437
|
+
const defaults = {};
|
|
438
|
+
if (j.defaults !== undefined) {
|
|
439
|
+
if (!j.defaults || typeof j.defaults !== 'object' || Array.isArray(j.defaults)) return null;
|
|
440
|
+
for (const [lane, d] of Object.entries(j.defaults)) {
|
|
441
|
+
if (!LANES.includes(lane)) return null;
|
|
442
|
+
if (!d || typeof d !== 'object' || Array.isArray(d)) return null;
|
|
443
|
+
const { model, effort, ...rest } = d;
|
|
444
|
+
if (Object.keys(rest).length) return null;
|
|
445
|
+
if (model !== undefined && badRouteValue('model', model)) return null;
|
|
446
|
+
if (effort !== undefined) {
|
|
447
|
+
if (badRouteValue('effort', effort)) return null;
|
|
448
|
+
if (!LANE_FLAGS[lane] || !LANE_FLAGS[lane].effort) return null; // a lane with no reasoning flag cannot have one pinned
|
|
449
|
+
}
|
|
450
|
+
defaults[lane] = { model: model ?? null, effort: effort ?? null };
|
|
451
|
+
}
|
|
452
|
+
}
|
|
453
|
+
return { enabled: j.enabled, defaults };
|
|
381
454
|
} catch {
|
|
382
455
|
return null;
|
|
383
456
|
}
|
|
384
457
|
}
|
|
385
458
|
|
|
459
|
+
// Kept as the narrow question most callers ask. null still means malformed.
|
|
460
|
+
export function enabledLanes(here = dirname(fileURLToPath(import.meta.url))) {
|
|
461
|
+
const c = laneConfig(here);
|
|
462
|
+
return c === null ? null : c.enabled;
|
|
463
|
+
}
|
|
464
|
+
|
|
465
|
+
// Flag beats lanes.json beats nothing. `source` is what makes the log audit-worthy:
|
|
466
|
+
// 'lane_default' means this run inherited the vendor CLI's own config, unseen from here.
|
|
467
|
+
export function resolveRoute(lane, opts, defaults) {
|
|
468
|
+
const d = (defaults && defaults[lane]) || {};
|
|
469
|
+
const model = opts.model ?? d.model ?? null;
|
|
470
|
+
const effort = opts.effort ?? d.effort ?? null;
|
|
471
|
+
const src = (flag, def) => (flag != null ? 'flag' : def != null ? 'lanes.json' : 'lane_default');
|
|
472
|
+
return { model, effort, model_source: src(opts.model, d.model), effort_source: src(opts.effort, d.effort) };
|
|
473
|
+
}
|
|
474
|
+
|
|
386
475
|
function usage(msg) {
|
|
387
476
|
if (msg) console.error('cli-run: ' + msg);
|
|
388
477
|
console.error(`usage: cli-run <${LANES.join('|')}> "<prompt>" [--brief FILE] [--timeout SECS] [--quiet]
|
|
389
|
-
[--expect-file PATH] [--expect-json]
|
|
478
|
+
[--model ID] [--effort LEVEL] [--expect-file PATH] [--expect-json]
|
|
390
479
|
cli-run codex --audit "<prompt>" read-only sandbox (audit shape)
|
|
391
|
-
cli-run qwen [--
|
|
392
|
-
cli-run --doctor [--run] enabled lanes, binaries
|
|
480
|
+
cli-run qwen [--safe-mode] "<prompt>" qwen-only flag
|
|
481
|
+
cli-run --doctor [--run] enabled lanes, binaries, and the route each one is pinned to
|
|
482
|
+
|
|
483
|
+
--model / --effort pin what a lane runs with, instead of letting it inherit its
|
|
484
|
+
own config. Every lane takes --model; every lane except qwen takes --effort.
|
|
485
|
+
Levels are the vendor's own (agy low|medium|high, hermes none|minimal|...): an
|
|
486
|
+
unknown level is rejected by the lane, and reported as that lane's exit code.
|
|
487
|
+
Pin them per lane instead of per call with "defaults" in bin/lanes.json.`);
|
|
393
488
|
return USAGE;
|
|
394
489
|
}
|
|
395
490
|
|
|
@@ -416,11 +511,12 @@ function installedPrimary(here = dirname(fileURLToPath(import.meta.url))) {
|
|
|
416
511
|
|
|
417
512
|
// --doctor: the first thing to run after install.
|
|
418
513
|
export async function doctor(run) {
|
|
419
|
-
const
|
|
420
|
-
if (
|
|
514
|
+
const cfg = laneConfig();
|
|
515
|
+
if (cfg === null) {
|
|
421
516
|
console.error('doctor: lanes.json exists but is malformed; fix it first');
|
|
422
517
|
return USAGE;
|
|
423
518
|
}
|
|
519
|
+
const { enabled, defaults } = cfg;
|
|
424
520
|
let bad = 0;
|
|
425
521
|
console.log(`doctor: ${enabled.length} enabled lane(s): ${enabled.join(', ') || 'none'}`);
|
|
426
522
|
const primary = installedPrimary();
|
|
@@ -432,7 +528,11 @@ export async function doctor(run) {
|
|
|
432
528
|
for (const lane of LANES) {
|
|
433
529
|
const on = enabled.includes(lane);
|
|
434
530
|
const bin = which(lane);
|
|
435
|
-
|
|
531
|
+
const d = defaults[lane] || {};
|
|
532
|
+
// A disabled lane has no route worth reporting; saying "not pinned" there
|
|
533
|
+
// reads as a finding about a lane that is not going to run.
|
|
534
|
+
const route = !on ? '' : d.model || d.effort ? `route ${d.model || 'lane default'}/${d.effort || 'lane default'}` : 'route not pinned (inherits the lane\'s own config)';
|
|
535
|
+
let line = ` ${lane.padEnd(7)} ${on ? 'enabled ' : 'disabled'} ${bin ? 'binary ok' : 'binary MISSING'}${route ? ' ' + route : ''}`;
|
|
436
536
|
if (on && !bin) bad++;
|
|
437
537
|
if (on && bin && run) {
|
|
438
538
|
const rc = await main([lane, 'Reply with exactly the word OK and nothing else.', '--timeout', '120', '--quiet']);
|
|
@@ -443,6 +543,7 @@ export async function doctor(run) {
|
|
|
443
543
|
}
|
|
444
544
|
console.log(bad ? `doctor: ${bad} problem(s)` : 'doctor: all enabled lanes ' + (run ? 'answered' : 'present'));
|
|
445
545
|
console.log('doctor checks presence and, with --run, a one-word canary. It does not check vendor versions.');
|
|
546
|
+
console.log('"route not pinned" means that lane runs on whatever its own config file says, which this tool cannot see. Pin it in lanes.json "defaults" if the route matters.');
|
|
446
547
|
return bad ? NO_DELIVERABLE : OK;
|
|
447
548
|
}
|
|
448
549
|
|
|
@@ -485,10 +586,10 @@ export function checkContracts(opts, text, before) {
|
|
|
485
586
|
}
|
|
486
587
|
|
|
487
588
|
export async function main(argv) {
|
|
488
|
-
const VALUE = new Set(['--brief', '--timeout', '--model', '--expect-file']);
|
|
589
|
+
const VALUE = new Set(['--brief', '--timeout', '--model', '--effort', '--expect-file']);
|
|
489
590
|
const BOOL = new Set(['--quiet', '--audit', '--safe-mode', '--doctor', '--run', '--expect-json']);
|
|
490
591
|
const args = [...argv];
|
|
491
|
-
const opts = { timeout: 900, quiet: false, audit: false, model: null, safeMode: false, brief: null, doctor: false, run: false, expectFile: null, expectJson: false };
|
|
592
|
+
const opts = { timeout: 900, quiet: false, audit: false, model: null, effort: null, safeMode: false, brief: null, doctor: false, run: false, expectFile: null, expectJson: false };
|
|
492
593
|
const positional = [];
|
|
493
594
|
while (args.length) {
|
|
494
595
|
const a = args.shift();
|
|
@@ -498,6 +599,7 @@ export async function main(argv) {
|
|
|
498
599
|
if (a === '--brief') opts.brief = v;
|
|
499
600
|
else if (a === '--timeout') opts.timeout = Number(v);
|
|
500
601
|
else if (a === '--expect-file') opts.expectFile = v;
|
|
602
|
+
else if (a === '--effort') opts.effort = v;
|
|
501
603
|
else opts.model = v;
|
|
502
604
|
} else if (BOOL.has(a)) {
|
|
503
605
|
if (a === '--quiet') opts.quiet = true;
|
|
@@ -530,11 +632,33 @@ export async function main(argv) {
|
|
|
530
632
|
if (!prompt) return usage('give a prompt or --brief FILE');
|
|
531
633
|
if (!Number.isFinite(opts.timeout) || opts.timeout <= 0) return usage('--timeout must be a positive number of seconds');
|
|
532
634
|
if (opts.audit && lane !== 'codex') return usage('--audit is codex-only');
|
|
533
|
-
if (
|
|
635
|
+
if (opts.safeMode && lane !== 'qwen') return usage('--safe-mode is qwen-only');
|
|
636
|
+
for (const [kind, v] of [['model', opts.model], ['effort', opts.effort]]) {
|
|
637
|
+
if (v == null) continue;
|
|
638
|
+
const bad = badRouteValue(kind, v);
|
|
639
|
+
if (bad) return usage(bad);
|
|
640
|
+
}
|
|
641
|
+
// qwen has no reasoning flag. Dropping --effort silently would leave the caller
|
|
642
|
+
// believing a route that never happened, which is the defect this feature fixes.
|
|
643
|
+
if (opts.effort && !(LANE_FLAGS[lane] && LANE_FLAGS[lane].effort)) return usage(`${lane} has no reasoning-effort flag; --effort is not available on this lane`);
|
|
534
644
|
|
|
535
645
|
const digest = createHash('sha256').update(prompt).digest('hex').slice(0, 12);
|
|
536
646
|
const base = { lane, prompt_sha256_12: digest, prompt_chars: prompt.length };
|
|
537
|
-
const
|
|
647
|
+
const cfg = laneConfig();
|
|
648
|
+
// Resolve the route BEFORE the refusals below. A run that never reached a lane
|
|
649
|
+
// was still a request for one, and a failure record with no route is the exact
|
|
650
|
+
// gap this feature exists to close. A malformed lanes.json has no usable
|
|
651
|
+
// defaults, so the flags stand alone and say so.
|
|
652
|
+
const route = resolveRoute(lane, opts, cfg === null ? {} : cfg.defaults);
|
|
653
|
+
opts.model = route.model;
|
|
654
|
+
opts.effort = route.effort;
|
|
655
|
+
Object.assign(base, {
|
|
656
|
+
model_requested: route.model,
|
|
657
|
+
effort_requested: route.effort,
|
|
658
|
+
model_source: route.model_source,
|
|
659
|
+
effort_source: route.effort_source
|
|
660
|
+
});
|
|
661
|
+
const enabled = cfg === null ? null : cfg.enabled;
|
|
538
662
|
if (enabled === null) {
|
|
539
663
|
console.error('cli-run: lanes.json exists but is not a valid {"enabled": [...]} file; refusing every lane until it is fixed');
|
|
540
664
|
log({ ...base, verdict: 'unavailable', rc: UNAVAILABLE, reason: 'lanes_json_malformed' });
|
|
@@ -601,7 +725,8 @@ export async function main(argv) {
|
|
|
601
725
|
}
|
|
602
726
|
}
|
|
603
727
|
if (text && code === OK) process.stdout.write(text + '\n');
|
|
604
|
-
|
|
728
|
+
const routeNote = route.model || route.effort ? `${route.model || 'lane default'}/${route.effort || 'lane default'}` : 'lane default';
|
|
729
|
+
if (!opts.quiet) console.error(`cli-run[${lane}] ${verdict} rc=${code} ${r.seconds.toFixed(1)}s raw=${r.outBytes || 0}B route=${routeNote} :: ${detail}`);
|
|
605
730
|
// Durable log: fixed reason code and structural numbers only.
|
|
606
731
|
log({ ...base, verdict, rc: code, cli_rc: r.status, signal: r.signal || null, seconds: Math.round(r.seconds * 100) / 100, raw_bytes: r.outBytes || 0, deliverable_bytes: Buffer.byteLength(text), reason: REASONS.has(reason) ? reason : 'unknown' });
|
|
607
732
|
return code;
|
package/bin/cli.js
CHANGED
|
@@ -93,7 +93,8 @@ Usage
|
|
|
93
93
|
Flags
|
|
94
94
|
--level 1|2|3 1 beginner (one agent), 2 intermediate (many CLIs), 3 advanced (plus a VM)
|
|
95
95
|
--ais a,b,c catalog ids you have access to (see --list)
|
|
96
|
-
--primary id the agent that runs the system and receives the subagents (any level
|
|
96
|
+
--primary id the agent that runs the system and receives the subagents (any level). When several qualify
|
|
97
|
+
and --yes is set, the run picks one and says so in the plan; pass this to decide it yourself.
|
|
97
98
|
--tools a,b companion tools to set up, all optional (default with --yes: codecalc only); --no-tools for none
|
|
98
99
|
--apis a,b level 3 only: metered API keys you HOLD (anthropic,openai,google,xai,openrouter); --no-apis for none.
|
|
99
100
|
Asked separately from the CLIs because a subscription is not an API key.
|
|
@@ -189,6 +190,7 @@ async function main() {
|
|
|
189
190
|
// 3. Primary agent (the one that runs the system)
|
|
190
191
|
const candidates = agentCandidates(selected);
|
|
191
192
|
let primary = null;
|
|
193
|
+
let primaryAutoPicked = false;
|
|
192
194
|
if (opt('primary')) {
|
|
193
195
|
primary = byId[opt('primary')];
|
|
194
196
|
if (!primary || !candidates.includes(primary)) bad('--primary must be one of: ' + candidates.map((a) => a.id).join(', '));
|
|
@@ -200,7 +202,10 @@ async function main() {
|
|
|
200
202
|
// --yes picks for the user: claude-code if present, else the first agent that can load subagent
|
|
201
203
|
// definitions (it gets five files written for it), else the first candidate. #19: codex listed
|
|
202
204
|
// before agy used to win and nothing was written to the project root.
|
|
203
|
-
if (yes)
|
|
205
|
+
if (yes) {
|
|
206
|
+
primary = candidates.find((a) => a.id === 'claude-code') || candidates.find((a) => a.agentsDir) || candidates[0];
|
|
207
|
+
primaryAutoPicked = true;
|
|
208
|
+
}
|
|
204
209
|
else {
|
|
205
210
|
console.log('\nWhich one is your primary agent (the one that runs the system)?');
|
|
206
211
|
candidates.forEach((a, i) => console.log(` ${i + 1} ${a.name}`));
|
|
@@ -271,7 +276,7 @@ async function main() {
|
|
|
271
276
|
const files = planFiles({ level, selected, primary, dir, project, tools, apis });
|
|
272
277
|
const lvl = LEVELS.find((l) => l.id === level);
|
|
273
278
|
const agentFiles = files.filter((f) => f.root === 'project');
|
|
274
|
-
console.log(`\nPlan\n level ${lvl.id} ${lvl.name}\n access ${selected.map((a) => a.id).join(', ')}\n primary ${primary ? primary.id : 'none'}\n tools ${tools.map((t) => t.id).join(', ') || 'none'}` + (level >= 3 ? `\n api keys ${apis.map((p) => p.id).join(', ') || 'none'}` : '') + `\n folder ${dir}\n project ${project}${agentFiles.length ? ' (' + agentFiles.length + ' subagent files go here)' : ''}\n files ${files.length}`);
|
|
279
|
+
console.log(`\nPlan\n level ${lvl.id} ${lvl.name}\n access ${selected.map((a) => a.id).join(', ')}\n primary ${primary ? primary.id : 'none'}${primaryAutoPicked ? ` (chosen for you from ${candidates.map((a) => a.id).join(', ')}; pass --primary to decide it yourself)` : ''}\n tools ${tools.map((t) => t.id).join(', ') || 'none'}` + (level >= 3 ? `\n api keys ${apis.map((p) => p.id).join(', ') || 'none'}` : '') + `\n folder ${dir}\n project ${project}${agentFiles.length ? ' (' + agentFiles.length + ' subagent files go here)' : ''}\n files ${files.length}`);
|
|
275
280
|
if (agentFiles.length && !opt('project')) {
|
|
276
281
|
console.log(`\nNote: --project was not given, so the ${agentFiles.length} subagent file(s) go to the current directory (${project}). Pass --project to put them somewhere else.`);
|
|
277
282
|
}
|
|
@@ -365,7 +370,7 @@ async function main() {
|
|
|
365
370
|
|
|
366
371
|
for (const t of tools) {
|
|
367
372
|
const doc = t.id.toUpperCase() + '.md';
|
|
368
|
-
console.log(`\n${t.name}\n
|
|
373
|
+
console.log(`\n${t.name}\n note: ${t.optionalNote}\n needs: ${t.requires}\n run: ${t.install}\n one-click or self-registering for: ${t.autoClients.join(', ')}. Other agents and the details: ${dir}/${doc}`);
|
|
369
374
|
}
|
|
370
375
|
|
|
371
376
|
// 7. Activation summary: writing the folder is half the job. Say exactly what
|
package/docs/part-1-beginner.md
CHANGED
|
@@ -14,6 +14,8 @@ If your agent exposes model choice (Claude Code, Codex, Antigravity), map the ti
|
|
|
14
14
|
|
|
15
15
|
Three cost levers, always together: tier (price per token), token discipline (how many tokens: read only what you will touch, never re-read, deliverables not narration), effort (how hard each call thinks).
|
|
16
16
|
|
|
17
|
+
And three inputs into the choice, not one. **Role** picks the agent. **Complexity** moves the effort: a worker executing a finished plan needs less reasoning than the reviewer judging its output, so when the plan is airtight the spec is carrying the thinking. **Risk** moves the tier and who reads the result: security, privacy, data loss and irreversible changes are the four worth naming, because none of their failures can be fixed by editing the code afterwards. A one-line change to an auth check is simple and high-risk at once, and it is the risk that decides.
|
|
18
|
+
|
|
17
19
|
Robustness first, cost second. You split tiers because the split produces better work.
|
|
18
20
|
|
|
19
21
|
## 2. Classify every task, first match wins
|
|
@@ -26,7 +26,9 @@ One driver, no second AI in the mix: the orchestrator invokes the CLIs; it never
|
|
|
26
26
|
|
|
27
27
|
Every agent CLI can report success and deliver nothing. `bin/cli-run.mjs` builds the right invocation per lane, reads that lane's **native** terminal event, and exits `10` when a run produced no deliverable, `12` on timeout, `13` when the lane is missing. Byte count is not a check either; a run can emit hundreds of kilobytes and no conclusion. One lane's own success flags lie outright (an upstream 400 reported as success), so its judge reads the two honest signals instead.
|
|
28
28
|
|
|
29
|
-
Every call goes through it. "This lane is flaky" becomes a query over its log instead of an argument. `node bin/cli-run.mjs --doctor` is the first thing to run after install: enabled lanes, binaries on PATH, and with `--run` a one-word canary per lane.
|
|
29
|
+
Every call goes through it. "This lane is flaky" becomes a query over its log instead of an argument. `node bin/cli-run.mjs --doctor` is the first thing to run after install: enabled lanes, binaries on PATH, the route each lane is pinned to, and with `--run` a one-word canary per lane.
|
|
30
|
+
|
|
31
|
+
There is a second thing a lane can be quietly wrong about. Left unpinned, it runs on **its own config file**, which the runner cannot see: a CLI set up months ago at a low reasoning effort keeps auditing at that effort while your routing docs describe an adversarial pass, and no error is ever raised. `--model` and `--effort` pin it per call, `defaults` in `bin/lanes.json` pins it per lane, and every run records the value requested and where it came from (`flag`, `lanes.json`, `lane_default`). The log claims no actual: grok reports a model id in its output, the other four lanes report none, so the field would be populated for one lane and empty for four, and it would be a provider-supplied string the durable log never holds.
|
|
30
32
|
|
|
31
33
|
## 4. Every delegation carries a task bundle, on both surfaces
|
|
32
34
|
|
|
@@ -36,6 +38,10 @@ Subagents and CLI lanes are the same problem: something with none of your rules
|
|
|
36
38
|
|
|
37
39
|
Fan the same plan to three model families (web sweep, adversarial read, live data), each as one `cli-run` call. The orchestrator opens the primary sources itself, marks every claim, and writes the only durable record. Expect one engine to return confident unsourced numerics; downgrade it. Weight the engines that report their own gaps. Count dispositions, not briefs.
|
|
38
40
|
|
|
41
|
+
## 5a. A finding is a claim, not a fact
|
|
42
|
+
|
|
43
|
+
An audit that returns six findings has returned six claims. Hand them to `finding-verifier` before any of them causes a repair: it reads the cited line, states what would trigger the problem, then hunts for the guard, caller or test that makes it impossible, and answers CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE. Only CONFIRMED earns a change. Use a different family from the one that produced the finding, and let INCONCLUSIVE stand: rounding it up to be safe buys unnecessary repairs, rounding it down to be tidy hides real ones.
|
|
44
|
+
|
|
39
45
|
## 6. Gap analysis gets a second family
|
|
40
46
|
|
|
41
47
|
The second pass is now a different model reading the same artifact, in read-only audit mode. Disagreement between families is the cheapest signal that something is soft.
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "model-orchestrator",
|
|
3
|
-
"version": "0.1.
|
|
4
|
-
"description": "Routing instructions and a CLI runner for your AI tools. One installer asks what you have access to and generates a matching setup, from one chat app to several agent CLIs or a virtual machine.",
|
|
3
|
+
"version": "0.1.14",
|
|
4
|
+
"description": "Routing instructions and a CLI runner for your AI tools. One installer asks what you have access to and generates a matching setup, from one chat app to several agent CLIs or a virtual machine. Routes by role, complexity and risk; pins and logs the model and reasoning effort each lane runs with.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
|
7
7
|
"model-orchestrator": "bin/cli.js"
|
|
@@ -51,7 +51,11 @@
|
|
|
51
51
|
"agents",
|
|
52
52
|
"subagents",
|
|
53
53
|
"qwen",
|
|
54
|
-
"mcp"
|
|
54
|
+
"mcp",
|
|
55
|
+
"reasoning-effort",
|
|
56
|
+
"model-routing",
|
|
57
|
+
"code-review",
|
|
58
|
+
"ai-agents"
|
|
55
59
|
],
|
|
56
60
|
"author": "aunysillyme (https://github.com/aunysillyme)",
|
|
57
61
|
"license": "MIT"
|
package/src/install.js
CHANGED
|
@@ -219,12 +219,32 @@ export function activationSteps(opts) {
|
|
|
219
219
|
else if (snippet) steps.push(`open ${primary.chatName || primary.name} and paste the block in ${join(dirAbs, snippet)} into its ${primary.chatSurface || 'custom instructions'}`);
|
|
220
220
|
if (primary && primary.agentsDir) steps.push(`subagents are in ${join(projectAbs, primary.agentsDir)}; run ${primary.bin} from ${projectAbs} to pick them up`);
|
|
221
221
|
for (const a of selected.filter((a) => a.bin && a.kind === 'agent-cli')) steps.push(`sign in to ${a.name}: ${a.auth}`);
|
|
222
|
+
// A local runtime has a bin but no sign-in, so the agent-cli loop above skips it
|
|
223
|
+
// and before this it appeared in no ordered list at any level (#26).
|
|
224
|
+
for (const a of selected.filter((a) => a.bin && a.kind === 'local')) steps.push(`install ${a.name}: ${a.install.url}, then \`${a.bin} pull <model>\` before the local lane can answer`);
|
|
222
225
|
for (const t of tools) steps.push(`${t.id}: ${t.install}`);
|
|
223
226
|
if (level >= 2) steps.push(`smoke test: node ${join(dirAbs, 'bin', 'cli-run.mjs')} --doctor (add --run to send each lane one tiny prompt)`);
|
|
224
227
|
if (level >= 3) steps.push(`box: read ${join(dirAbs, 'vm', 'README.md')}; keys named in vm/ENVIRONMENT.md go in your secrets manager, never a file`);
|
|
225
228
|
return steps;
|
|
226
229
|
}
|
|
227
230
|
|
|
231
|
+
// The verification list, in order. Gated on level for the same reason
|
|
232
|
+
// activationSteps is: level 1 writes no bin/, so a step naming cli-run.mjs or
|
|
233
|
+
// lanes.json there described an install that did not happen (#27).
|
|
234
|
+
export function proofSteps(opts) {
|
|
235
|
+
const { level } = opts;
|
|
236
|
+
const steps = [
|
|
237
|
+
'Start a fresh agent session and ask: "Read the orchestrator instructions. Quote the routing rule you will use, then sort pear, apple, banana alphabetically. Name the tier and whether you delegated."',
|
|
238
|
+
'Expect the fast tier and `apple, banana, pear`. If the agent cannot quote the routing rule, check the snippet location or chat instructions before continuing. This is a manual activation check, not proof that every future task follows the rules.'
|
|
239
|
+
];
|
|
240
|
+
if (level >= 2) {
|
|
241
|
+
steps.push('Run `node bin/cli-run.mjs --doctor` from this folder. It checks binary presence, not authentication or loaded instructions, and prints the model and effort each lane is pinned to. `--doctor --run` additionally uses a little quota to test live responses. No enabled lanes means delegation is inactive.');
|
|
242
|
+
steps.push('Decide whether the route matters to you. Every lane starts unpinned, which means it runs on whatever its own config file says: a CLI configured months ago at a low reasoning effort will keep auditing at that effort while your docs describe something stronger. Pin it in `bin/lanes.json` under `defaults`, or per call with `--model` and `--effort`. Either way the run is recorded in the log with the value requested and where it came from.');
|
|
243
|
+
steps.push('To test a real output contract, choose an enabled lane from `bin/lanes.json` and run `node bin/cli-run.mjs <lane> \'Return only {"sorted":["apple","banana","pear"]}\' --expect-json`. This uses quota. Expect JSON and exit 0; inspect the array yourself. A non-JSON response exits 10, a missing binary exits 13, and an authentication failure reports the vendor error. The explicit lane tests execution; your primary agent still makes delegation decisions.');
|
|
244
|
+
}
|
|
245
|
+
return steps;
|
|
246
|
+
}
|
|
247
|
+
|
|
228
248
|
function vars(opts) {
|
|
229
249
|
const { level, selected, primary } = opts;
|
|
230
250
|
const tools = opts.tools || [];
|
|
@@ -240,6 +260,7 @@ function vars(opts) {
|
|
|
240
260
|
const pinOf = (id) => (toolById[id] && toolById[id].pin) || 'latest';
|
|
241
261
|
const snippet = snippetFor(primary);
|
|
242
262
|
const steps = activationSteps({ level, selected, primary, tools, dir: opts.dir, project: opts.project });
|
|
263
|
+
const proofs = proofSteps({ level });
|
|
243
264
|
// Only claude-code and agy put files under the project root. A chat primary
|
|
244
265
|
// puts nothing there, so naming a project root would name a folder this run
|
|
245
266
|
// never created (#21).
|
|
@@ -253,6 +274,7 @@ function vars(opts) {
|
|
|
253
274
|
return {
|
|
254
275
|
...laneVars(selected),
|
|
255
276
|
ACTIVATION_STEPS: steps.map((st, i) => `${i + 1}. ${st}`).join('\n'),
|
|
277
|
+
PROOF_STEPS: proofs.map((st, i) => `${i + 1}. ${st}`).join('\n'),
|
|
256
278
|
LOAD_IT: readsProjectRules
|
|
257
279
|
? `${primary.name} reads its rules from \`${primary.rulesFile}\` in the project root. The installer wrote \`${snippet}\` next to this README; copy its contents into \`${join(projectAbs, primary.rulesFile)}\`, creating that file if it does not exist. Nothing was appended to a file you already had.`
|
|
258
280
|
: snippet
|
|
@@ -272,6 +294,12 @@ function vars(opts) {
|
|
|
272
294
|
INSTALL_DIR: dirAbs,
|
|
273
295
|
INSTALL_DIR_SH: shellQuote(dirAbs),
|
|
274
296
|
INSTALL_DIR_SYSTEMD: systemdEscape(dirAbs),
|
|
297
|
+
// vm/README.md step 3 named `grok login` and `agy` whatever you picked (#26).
|
|
298
|
+
VM_SIGNIN: (() => {
|
|
299
|
+
const lines = selected.filter((a) => a.bin && a.kind === 'agent-cli').map((a) => ` - ${a.name}: ${a.auth}`);
|
|
300
|
+
for (const a of selected.filter((a) => a.bin && a.kind === 'local')) lines.push(` - ${a.name}: no sign-in. Install it from ${a.install.url}, then \`${a.bin} pull <model>\`.`);
|
|
301
|
+
return lines.length ? lines.join('\n') : ' - none: no CLI you selected needs a sign-in on the box.';
|
|
302
|
+
})(),
|
|
275
303
|
AUDIT_LANE: lane || 'none',
|
|
276
304
|
// Enforced boundary per lane: codex has a read-only sandbox flag; the others
|
|
277
305
|
// run with whatever their own config allows, and the script says so.
|
|
@@ -362,7 +390,9 @@ export function planFiles(opts) {
|
|
|
362
390
|
JSON.stringify(
|
|
363
391
|
{
|
|
364
392
|
enabled: selected.filter((a) => a.cliRun).map((a) => a.id),
|
|
365
|
-
|
|
393
|
+
defaults: {},
|
|
394
|
+
note: 'Lanes cli-run may call. Edit to enable or disable a lane. A lane not listed here exits 13 (unavailable).',
|
|
395
|
+
defaultsNote: 'Pin what a lane runs with, so the route in your docs is the route that runs: "defaults": {"codex": {"model": "gpt-6-astra", "effort": "high"}}. Left empty, a lane inherits its own config file, which cli-run cannot see and does not guess. `--model` and `--effort` override this per call, and `--doctor` prints what each lane is pinned to. Every lane takes a model; every lane except qwen takes an effort.'
|
|
366
396
|
},
|
|
367
397
|
null,
|
|
368
398
|
2
|
|
@@ -23,7 +23,8 @@ Generated {{DATE}} for: `{{AI_IDS}}`. Installed at `{{INSTALL_DIR}}`; the system
|
|
|
23
23
|
|
|
24
24
|
1. Provision a box. Ubuntu, 2+ vCPU, 8 GB is comfortable. Put it on a private mesh network if you can; do not open ports to the internet.
|
|
25
25
|
2. `bash setup-vm.sh`. It installs system deps and the npm-installable CLIs, then **prints** the vendor shell installers for the rest. Read those scripts before running them.
|
|
26
|
-
3. Sign each CLI in
|
|
26
|
+
3. Sign each CLI in, using the flow its vendor gives you. Run these inside `tmux` so a dropped SSH session does not kill the prompt. Headless Linux has no keyring by default; `setup-vm.sh` installs one so the CLIs stop re-prompting.
|
|
27
|
+
{{VM_SIGNIN}}
|
|
27
28
|
4. Put provider keys in your secrets manager and export the names listed in `ENVIRONMENT.md` into the gateway's environment at start time. Never write a value into a file in this folder. The gateway config was rendered from the API keys you said you hold, not from your CLI subscriptions: those are different entitlements.
|
|
28
29
|
5. `docker compose up -d`, then list the lanes without putting the key in argv (the key must be a single token, `^[A-Za-z0-9._-]+$`, because it is interpolated into curl's config grammar):
|
|
29
30
|
```bash
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
# .agents/agents/
|
|
2
2
|
|
|
3
|
-
Antigravity CLI custom agents, one per tier, in the `.agents/agents/<name>.md` format (YAML frontmatter + system prompt). `model` is a tier (`flash`, `pro`) or `inherit`. `subagent: true` lets a coordinator call them through `invoke_subagent`, which takes an array and launches concurrently; `mainAgent: true` lets you launch them directly with `agy --agent <name>`.
|
|
3
|
+
Antigravity CLI custom agents, one per tier plus a finding-verifier, in the `.agents/agents/<name>.md` format (YAML frontmatter + system prompt). `model` is a tier (`flash`, `pro`) or `inherit`. `subagent: true` lets a coordinator call them through `invoke_subagent`, which takes an array and launches concurrently; `mainAgent: true` lets you launch them directly with `agy --agent <name>`.
|
|
4
4
|
|
|
5
5
|
`commandExecutionPolicy` is `auto` for `builder` (it has to run builds and tests; `auto` keeps deletes and other high-risk commands gated) and `off` for the read-only agents. `model` is a tier: `pro` for deep-planner, `flash` for the rest.
|
|
@@ -0,0 +1,29 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: finding-verifier
|
|
3
|
+
description: Adversarial verification of review findings; tries to disprove each one and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE. Read-only, never repairs.
|
|
4
|
+
model: flash
|
|
5
|
+
subagent: true
|
|
6
|
+
mainAgent: true
|
|
7
|
+
commandExecutionPolicy: off
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
# finding-verifier
|
|
11
|
+
|
|
12
|
+
A finding is a claim, not a fact. You try to disprove each one before it is
|
|
13
|
+
allowed to cause a repair.
|
|
14
|
+
|
|
15
|
+
For each finding you are given: read the cited file and line yourself, state the
|
|
16
|
+
input or sequence that would trigger it, then hunt for what makes it impossible
|
|
17
|
+
(a guard upstream, a caller that never passes that value, an existing test).
|
|
18
|
+
|
|
19
|
+
Return one verdict per finding, in the order given:
|
|
20
|
+
- CONFIRMED: reproduced, or a concrete unblocked path. Give the path.
|
|
21
|
+
- NOT_REPRODUCED: you found what stops it. Name it and where it is.
|
|
22
|
+
- INCONCLUSIVE: not settleable read-only. Say what you would need.
|
|
23
|
+
|
|
24
|
+
Rules:
|
|
25
|
+
- Stay inside the task bundle you were given. Anything not granted is denied.
|
|
26
|
+
- Verify only the findings handed to you; anything else you notice goes at the end, marked unverified.
|
|
27
|
+
- Never round INCONCLUSIVE up to CONFIRMED to be safe, or down to NOT_REPRODUCED to be tidy.
|
|
28
|
+
- Read-only: you never repair and never reword a finding.
|
|
29
|
+
- Token discipline: read the cited code and its callers, not the repository.
|
|
@@ -1,12 +1,13 @@
|
|
|
1
1
|
# .claude/agents/
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
Six subagents. Five are one per tier; `finding-verifier` is the check that sits between a review and a repair. Claude Code loads project-level agents from this folder automatically.
|
|
4
4
|
|
|
5
5
|
| Agent | Tier | Model alias | Effort | Job |
|
|
6
6
|
|---|---|---|---|---|
|
|
7
7
|
| deep-planner | deep | opus | xhigh | judges every build twice; never retrieves |
|
|
8
8
|
| builder | standard | sonnet | high | bounded sub-parts of a build |
|
|
9
9
|
| code-reviewer | standard | sonnet | high | read-only findings |
|
|
10
|
+
| finding-verifier | standard | sonnet | high | tries to disprove a finding before it causes a repair |
|
|
10
11
|
| live-researcher | standard | sonnet | medium | fresh data through tools |
|
|
11
12
|
| bulk-worker | fast | haiku | low | mechanical volume |
|
|
12
13
|
|
|
@@ -0,0 +1,43 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: finding-verifier
|
|
3
|
+
description: Adversarial verification of review findings. Use after a review or audit returns findings and before any of them trigger a repair. Read-only. Tries to DISPROVE each finding and returns CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Do not use to find new problems, and do not use to fix anything.
|
|
4
|
+
tools: Read, Glob, Grep, Bash
|
|
5
|
+
model: sonnet
|
|
6
|
+
effort: high
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
You are the verification tier of the model router.
|
|
10
|
+
|
|
11
|
+
A finding is a claim, not a fact. Your job is to try to disprove each one before
|
|
12
|
+
it is allowed to cause a change. A false finding is expensive twice: it buys a
|
|
13
|
+
repair nobody needed, and it teaches everyone to skim the next report.
|
|
14
|
+
|
|
15
|
+
You are given findings from a review or an audit. For each one, independently:
|
|
16
|
+
|
|
17
|
+
1. Read the cited file and line yourself. A citation that does not point at what
|
|
18
|
+
the finding describes is already a failure of the finding, not of the code.
|
|
19
|
+
2. State the exact input, state or sequence that would make it happen.
|
|
20
|
+
3. Look for what makes it impossible: a guard upstream, a type that cannot hold
|
|
21
|
+
that value, a caller that never passes it, a test that already covers it, a
|
|
22
|
+
framework guarantee.
|
|
23
|
+
4. Where you can run something cheap and read-only that settles it, run it.
|
|
24
|
+
|
|
25
|
+
Return one verdict per finding, in the order you were given them:
|
|
26
|
+
|
|
27
|
+
- **CONFIRMED** you reproduced it, or traced a concrete path to it that nothing
|
|
28
|
+
prevents. Give the path in one or two sentences.
|
|
29
|
+
- **NOT_REPRODUCED** you found what stops it. Name that thing and where it is.
|
|
30
|
+
This is a success, not a failure to try.
|
|
31
|
+
- **INCONCLUSIVE** you could not settle it read-only. Say exactly what you would
|
|
32
|
+
need: a test run, a credential, a live environment, a decision from a human.
|
|
33
|
+
Never round this up to CONFIRMED to be safe, and never down to
|
|
34
|
+
NOT_REPRODUCED to be tidy.
|
|
35
|
+
|
|
36
|
+
Rules:
|
|
37
|
+
- Verify only the findings you were given. New problems you happen to notice go
|
|
38
|
+
in a separate list at the end, clearly marked as unverified observations.
|
|
39
|
+
- You are read-only. You never repair, and you never soften a finding's wording.
|
|
40
|
+
- Verifying nothing is a real answer. If every finding is NOT_REPRODUCED, say
|
|
41
|
+
that plainly; a verifier that always confirms something is a rubber stamp
|
|
42
|
+
facing the other way.
|
|
43
|
+
- Token discipline: read the cited code and its callers, not the repository.
|
|
@@ -10,6 +10,7 @@ Routing rules live in `{{RULES_PATH}}/{{ROUTING_FILE}}`. Read them before any bu
|
|
|
10
10
|
1. Bulk, mechanical, many similar items -> bulk-worker (fast tier).
|
|
11
11
|
2. Needs live data -> live-researcher (standard tier + tools).
|
|
12
12
|
3. Review without changing -> code-reviewer (standard, read-only).
|
|
13
|
+
3a. Holding findings from a review or scanner -> finding-verifier before any repair. Only CONFIRMED findings earn a change.
|
|
13
14
|
4. Ambiguous, architectural, or expensive to get wrong -> deep-planner (deep tier), then hand the plan down.
|
|
14
15
|
5. Everything else that changes files -> build it directly. The main build is never handed off whole; bounded sub-parts go to builder.
|
|
15
16
|
|
|
@@ -21,10 +21,7 @@ These are the same steps, in the same order, that the installer printed in your
|
|
|
21
21
|
|
|
22
22
|
## Then prove it took
|
|
23
23
|
|
|
24
|
-
|
|
25
|
-
2. Expect the fast tier and `apple, banana, pear`. If the agent cannot quote the routing rule, check the snippet location or chat instructions before continuing. This is a manual activation check, not proof that every future task follows the rules.
|
|
26
|
-
3. At level 2+, run `node bin/cli-run.mjs --doctor` from this folder. It checks binary presence, not authentication or loaded instructions. `--doctor --run` additionally uses a little quota to test live responses. No enabled lanes means delegation is inactive.
|
|
27
|
-
4. To test a real output contract, choose an enabled lane from `bin/lanes.json` and run `node bin/cli-run.mjs <lane> 'Return only {"sorted":["apple","banana","pear"]}' --expect-json`. This uses quota. Expect JSON and exit 0; inspect the array yourself. A non-JSON response exits 10, a missing binary exits 13, and an authentication failure reports the vendor error. The explicit lane tests execution; your primary agent still makes delegation decisions.
|
|
24
|
+
{{PROOF_STEPS}}
|
|
28
25
|
|
|
29
26
|
## What is in this folder
|
|
30
27
|
|
|
@@ -66,7 +66,17 @@ Secret detection, static analysis and dependency scanning, filtered to lines thi
|
|
|
66
66
|
|
|
67
67
|
Route by shape: security-shaped diffs (auth, tokens, routes, deletion, bulk mutation, untrusted input) go to an adversarial auditor, ideally a **different model family**. Architecture-shaped diffs go to the judgment tier reviewing build against plan. Never both on one diff.
|
|
68
68
|
|
|
69
|
-
|
|
69
|
+
Findings do not go straight to a repair. Hand them to the **finding-verifier**, whose job is to DISPROVE each one: read the cited line, state what would trigger it, then hunt for the guard, caller or test that makes it impossible. It returns one of three verdicts per finding, and rounding between them is the failure mode to watch for.
|
|
70
|
+
|
|
71
|
+
| Verdict | Meaning | What happens next |
|
|
72
|
+
|---|---|---|
|
|
73
|
+
| CONFIRMED | reproduced, or a concrete path nothing blocks | it earns a repair |
|
|
74
|
+
| NOT_REPRODUCED | something prevents it, named and located | dropped, and not narrated |
|
|
75
|
+
| INCONCLUSIVE | not settleable read-only | say what it would take; never round it to either side |
|
|
76
|
+
|
|
77
|
+
Use a different model family from the one that produced the finding where you have one: a family asked to check its own claim tends to agree with itself.
|
|
78
|
+
|
|
79
|
+
**Gate:** every finding **verified** before it reaches a human, and only CONFIRMED findings trigger a change. Hard cap one re-audit. `CLEAN` is a valid success state; an auditor that is not allowed to say so manufactures something, and so does a verifier that is expected to confirm.
|
|
70
80
|
|
|
71
81
|
### Stage 5b · Ship gate
|
|
72
82
|
1. What is the rollback target? Record it before shipping.
|
|
@@ -18,7 +18,8 @@ Enabled lanes (edit `bin/lanes.json`): {{CLI_RUN_LANES}}
|
|
|
18
18
|
```bash
|
|
19
19
|
node bin/cli-run.mjs <grok|codex|agy|hermes|qwen> "<prompt>" [--brief FILE] [--timeout SECS] [--quiet]
|
|
20
20
|
node bin/cli-run.mjs codex --audit "<prompt>" # read-only sandbox, the audit shape
|
|
21
|
-
node bin/cli-run.mjs
|
|
21
|
+
node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
|
|
22
|
+
node bin/cli-run.mjs qwen [--safe-mode] "<prompt>" # qwen-only flag
|
|
22
23
|
```
|
|
23
24
|
|
|
24
25
|
Put it on your PATH if you like: `ln -s "$PWD/bin/cli-run.mjs" ~/.local/bin/cli-run`.
|
|
@@ -65,13 +66,49 @@ qwen is the lane whose own success flags lie: an upstream 400 comes back as exit
|
|
|
65
66
|
| 2 | usage error in cli-run itself |
|
|
66
67
|
| N | the lane exited N != 0: passed through unchanged, verdict `exit_nonzero`, even when parseable text came back. The bounded head of the lane's stderr is shown on your terminal so an auth failure reads as one |
|
|
67
68
|
|
|
69
|
+
## The route: which model, and how hard it thinks
|
|
70
|
+
|
|
71
|
+
A lane you do not pin runs on **its own config file**, which this tool cannot see. That is the quiet failure this section exists for: a CLI configured months ago at `reasoning_effort = "low"` keeps auditing at low effort while your routing docs describe an adversarial pass, and nothing anywhere says so.
|
|
72
|
+
|
|
73
|
+
Pin it per call, or per lane:
|
|
74
|
+
|
|
75
|
+
```bash
|
|
76
|
+
node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high # this call only
|
|
77
|
+
node bin/cli-run.mjs --doctor # prints what each lane is pinned to
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
```json
|
|
81
|
+
{
|
|
82
|
+
"enabled": ["codex", "grok"],
|
|
83
|
+
"defaults": { "codex": { "model": "gpt-6-astra", "effort": "high" } }
|
|
84
|
+
}
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
A flag beats `defaults`; `defaults` beats nothing. Each vendor spells these differently and `cli-run` translates:
|
|
88
|
+
|
|
89
|
+
| Lane | Model | Reasoning effort |
|
|
90
|
+
|---|---|---|
|
|
91
|
+
| grok | `-m` | `--reasoning-effort` |
|
|
92
|
+
| codex | `-m` | `-c model_reasoning_effort="LEVEL"` |
|
|
93
|
+
| agy | `--model` | `--effort` (low, medium, high) |
|
|
94
|
+
| hermes | `-m` | `--reasoning` (none, minimal, ...) |
|
|
95
|
+
| qwen | `-m` | none: this lane has no reasoning flag |
|
|
96
|
+
|
|
97
|
+
Three rules that keep this honest:
|
|
98
|
+
|
|
99
|
+
- **A level `cli-run` does not recognise is not rejected here.** Levels are the vendor's, they change, and guessing the valid set would date this tool. An unknown level is refused by the lane and surfaces as that lane's own exit code and stderr.
|
|
100
|
+
- **`--effort` on qwen is a usage error, not a silent drop.** A flag that vanishes leaves you believing a route that never ran.
|
|
101
|
+
- **Values are charset-bounded** (letters, digits, and `. _ : @ / + -`, no leading dash, 64 characters). A model id becomes an argv element and, on codex, part of a TOML value; bounding it is what stops either from being escaped.
|
|
102
|
+
|
|
68
103
|
## Permissions are a separate layer
|
|
69
104
|
|
|
70
105
|
`cli-run` never injects permission flags. Each CLI carries its own config, so every caller gets the same behaviour. Use each vendor's deny-list as the base layer; allow-lists only hold if every binary is enumerable in advance.
|
|
71
106
|
|
|
72
107
|
## Log
|
|
73
108
|
|
|
74
|
-
`~/.ai-orchestrator/cli-run.log.jsonl`, one line per run: lane, verdict, rc, the lane's own exit code, signal, seconds, raw bytes, deliverable bytes, a 12-hex sha256 prefix of the prompt and its length, and `reason`: one of a fixed set of codes (`ok`, `not_json`, `bad_stop_reason`, `empty_text`, `no_terminal_event`, `bad_status`, `api_error_in_result`, `total_errors`, `contract_unmet`, `exit_nonzero`, `timeout`, `killed`, `disabled`, `lanes_json_malformed`, ...). Never the prompt text, never a provider-supplied value, never free text: a value the log does not recognise is written as `unknown`. The human-readable detail, which may quote the provider, goes to your terminal only (and nowhere with `--quiet`). "This lane is flaky" becomes a query instead of an argument.
|
|
109
|
+
`~/.ai-orchestrator/cli-run.log.jsonl`, one line per run: lane, verdict, rc, the lane's own exit code, signal, seconds, raw bytes, deliverable bytes, a 12-hex sha256 prefix of the prompt and its length, the route (`model_requested`, `effort_requested`, and `model_source` / `effort_source`, each one of `flag`, `lanes.json` or `lane_default`), and `reason`: one of a fixed set of codes (`ok`, `not_json`, `bad_stop_reason`, `empty_text`, `no_terminal_event`, `bad_status`, `api_error_in_result`, `total_errors`, `contract_unmet`, `exit_nonzero`, `timeout`, `killed`, `disabled`, `lanes_json_malformed`, ...). Never the prompt text, never a provider-supplied value, never free text: a value the log does not recognise is written as `unknown`. The human-readable detail, which may quote the provider, goes to your terminal only (and nowhere with `--quiet`). "This lane is flaky" becomes a query instead of an argument, and so does "we route audits at high effort".
|
|
110
|
+
|
|
111
|
+
The log records what was **requested**, on every record including a run refused before the lane started. It does not record an actual. Reporting is inconsistent: grok returns a `modelUsage` block naming a model, the other four lanes return nothing of the kind, so an `actual` field would be populated for one lane and empty for four. It would also be a provider-supplied string, and this log holds fixed codes and bounded caller-supplied values only. `model_source: "lane_default"` is the honest way to say this run inherited something invisible from here.
|
|
75
112
|
|
|
76
113
|
## The prompt travels in argv
|
|
77
114
|
|
|
@@ -79,7 +116,7 @@ That is each vendor's documented headless shape (`-p`, `exec`). Two consequences
|
|
|
79
116
|
|
|
80
117
|
## lanes.json fails closed
|
|
81
118
|
|
|
82
|
-
Absent: every lane enabled. Present but malformed or unreadable: every lane refused (exit 13) until it is fixed. A half-written config never re-enables a lane the installer disabled.
|
|
119
|
+
Absent: every lane enabled, nothing pinned. Present but malformed or unreadable: every lane refused (exit 13) until it is fixed. A half-written config never re-enables a lane the installer disabled. `defaults` is optional and held to the same standard: a malformed entry, an unknown lane, an unknown key, a value outside the charset, or an effort pinned on a lane that has no reasoning flag all fail the whole file closed rather than being skipped quietly.
|
|
83
120
|
|
|
84
121
|
## A killed lane is not a deliverable
|
|
85
122
|
|
|
@@ -17,6 +17,7 @@ Rule of thumb: never spend a frontier token on a task a cheap tier finishes corr
|
|
|
17
17
|
1. **Bulk and mechanical?** → fast tier{{BULK_LANE}}. Many independent items each needing its own agent turn → a concurrent fan-out lane if you have one.
|
|
18
18
|
2. **Needs live data?** → {{LIVE_LANE}} standard tier with web tools.
|
|
19
19
|
3. **Reviewing without changing?** → standard tier read-only. Security-critical → {{ATTACK_LANE}}.
|
|
20
|
+
3a. **Holding findings from a review or a scanner?** → finding-verifier before any of them cause a repair. A finding is a claim, not a fact.
|
|
20
21
|
4. **Ambiguous, strategic, expensive to get wrong?** → deep tier (deep-planner). Then hand the plan down.
|
|
21
22
|
5. **Everything else that changes files** → the orchestrator builds it directly. Bounded sub-parts go to cheaper tiers; the main build is never handed off whole.
|
|
22
23
|
|
|
@@ -38,6 +39,7 @@ Every delegation carries `TASK_BUNDLE.md`. Its brief must restate every conventi
|
|
|
38
39
|
| 3 Build | the orchestrator, against the installed dependency's source |
|
|
39
40
|
| 4 Scan | secret + static + dependency scanners, diff-scoped, fail closed |
|
|
40
41
|
| 5 Attack | security-shaped diff → {{ATTACK_LANE}}. Architecture-shaped → deep tier, build against plan. Never both |
|
|
42
|
+
| 5a Verify findings | finding-verifier, a different model family where you have one: CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Only CONFIRMED earns a repair |
|
|
41
43
|
| 5b Ship | rollback id recorded, explicit human yes |
|
|
42
44
|
| 6 Verify | real test, negative test seen red, old identifier re-grepped to zero |
|
|
43
45
|
| 7 Record | one end-to-end doc, tracker Done with evidence, plan doc deleted |
|
|
@@ -59,7 +61,9 @@ One writer per run; every other lane proposes. Search before writing, index in t
|
|
|
59
61
|
- **De-escalation:** a request that sounds deep but is a lookup routes down.
|
|
60
62
|
- **Long context:** mechanical digestion → fast tier in chunks; judgment over a long input → standard tier.
|
|
61
63
|
- **Token discipline on every delegation:** pass only the context the delegate needs, never the conversation.
|
|
62
|
-
- **Effort per agent:** deep xhigh, review and build high, live research medium, bulk low.
|
|
64
|
+
- **Effort per agent:** deep xhigh, review, verification and build high, live research medium, bulk low.
|
|
65
|
+
- **Three inputs, not one:** role picks the agent, complexity moves the effort, risk moves the tier and who reads it. A one-line auth change is simple and high-risk at once, and the risk decides. See `TIERS.md`.
|
|
66
|
+
- **Pin the route when it matters:** a lane with no `--model`/`--effort` and no `defaults` entry in `bin/lanes.json` runs on its own config, which may be nothing like what this file describes. `cli-run --doctor` prints what each lane is pinned to, and every run logs the value requested and where it came from.
|
|
63
67
|
|
|
64
68
|
## Example routings
|
|
65
69
|
|
|
@@ -70,4 +74,5 @@ One writer per run; every other lane proposes. Search before writing, index in t
|
|
|
70
74
|
| "Add an endpoint" | the orchestrator builds it |
|
|
71
75
|
| "Why does this silently drop rows sometimes" | deep-planner (unknown cause), then build the fix directly |
|
|
72
76
|
| "Summarize these 30 notes into one index" | bulk-worker |
|
|
77
|
+
| "The audit returned 6 findings" | finding-verifier first; repair only what comes back CONFIRMED |
|
|
73
78
|
{{LANE_EXAMPLES}}
|
|
@@ -19,11 +19,52 @@ Tier sets the price per token. Token discipline sets how many tokens. **Effort s
|
|
|
19
19
|
|---|---|---|---|
|
|
20
20
|
| deep-planner | deep | xhigh | judges every build twice; expensive to get wrong |
|
|
21
21
|
| code-reviewer | standard | high | every endpoint is internet-facing |
|
|
22
|
+
| finding-verifier | standard | high | judging a claim is harder than producing it |
|
|
22
23
|
| builder | standard | high | a botched deploy is the costly failure |
|
|
23
24
|
| live-researcher | standard | medium | tools do the retrieval |
|
|
24
25
|
| bulk-worker | fast | low | the biggest cost win |
|
|
25
26
|
|
|
26
|
-
|
|
27
|
+
## Three inputs, not one
|
|
28
|
+
|
|
29
|
+
Role alone does not decide a route. Two more inputs move it, and they move it in
|
|
30
|
+
opposite directions, so state them separately instead of folding them into the
|
|
31
|
+
role.
|
|
32
|
+
|
|
33
|
+
**Complexity moves the effort.** The same role does not need the same reasoning
|
|
34
|
+
on every task.
|
|
35
|
+
|
|
36
|
+
| Complexity | What it looks like | What moves |
|
|
37
|
+
|---|---|---|
|
|
38
|
+
| simple | one file, one obvious edit, no unknowns | drop one effort level |
|
|
39
|
+
| standard | the default | the table above |
|
|
40
|
+
| complex | several surfaces, or an unknown cause | keep effort, add the deep-tier checkpoint |
|
|
41
|
+
| critical | irreversible, or it rewrites a standing rule | the escalation rule below applies |
|
|
42
|
+
|
|
43
|
+
The dial that pays for itself: **a worker executing a finished plan needs less
|
|
44
|
+
reasoning than the reviewer judging its output.** When the plan is airtight the
|
|
45
|
+
spec is carrying the thinking, so builder drops to medium. When the plan is
|
|
46
|
+
vague, fix the plan; do not buy reasoning to paper over it.
|
|
47
|
+
|
|
48
|
+
**Risk moves the tier and the reader, never just the effort.** These four are
|
|
49
|
+
the ones worth naming, because their failures are not recoverable by editing the
|
|
50
|
+
code afterwards.
|
|
51
|
+
|
|
52
|
+
| Risk | Present when the change touches | What it buys |
|
|
53
|
+
|---|---|---|
|
|
54
|
+
| security | auth, tokens, sessions, routes, untrusted input | the attack pass, ideally a different model family |
|
|
55
|
+
| privacy | personal data, anything leaving the machine | the local lane, and a named check on what is sent |
|
|
56
|
+
| data loss | deletion, bulk mutation, migrations, overwrites | a reviewed rollback path before the change is written |
|
|
57
|
+
| irreversible | publishing, sending, rotating, anything with an audience | a human yes at Stage 5b, never an agent's |
|
|
58
|
+
|
|
59
|
+
A risk raises code-reviewer to xhigh, and a security-shaped diff goes to the
|
|
60
|
+
attack lane rather than to a second read by the same family. Risk is not a
|
|
61
|
+
synonym for difficulty: a one-line change to an auth check is simple and
|
|
62
|
+
high-risk at the same time, and it is the risk that decides the route.
|
|
63
|
+
|
|
64
|
+
**Reserve the top of the ladder for evidence.** xhigh and the escalation tier are
|
|
65
|
+
bought with a named reason: a reproduced failure, a checkpoint that came back
|
|
66
|
+
unresolved, an irreversible change. A task that merely feels hard is a deep-tier
|
|
67
|
+
task, not an escalation.
|
|
27
68
|
|
|
28
69
|
## Why split tiers: robustness first, cost second
|
|
29
70
|
|