model-orchestrator 0.1.35 → 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (108) hide show
  1. package/AGENTS.md +31 -21
  2. package/CHANGELOG.md +43 -1
  3. package/README.md +127 -110
  4. package/bin/README.md +57 -6
  5. package/bin/aunx.js +7 -0
  6. package/bin/cli-run.mjs +21 -15
  7. package/bin/cli.js +376 -257
  8. package/docs/README.md +15 -18
  9. package/docs/catalog.md +228 -38
  10. package/docs/companions.md +28 -10
  11. package/docs/guarantees.md +21 -12
  12. package/docs/how-it-routes.md +49 -42
  13. package/docs/install.md +135 -33
  14. package/docs/part-1-beginner.md +37 -45
  15. package/docs/part-2-intermediate.md +34 -52
  16. package/docs/part-3-advanced.md +36 -26
  17. package/docs/security-review-history.md +38 -0
  18. package/llms.txt +24 -25
  19. package/package.json +15 -8
  20. package/proof/README.md +100 -0
  21. package/proof/gate-demo.cast +9 -0
  22. package/proof/gate-demo.gif +0 -0
  23. package/proof/results.json +198 -0
  24. package/proof/scripts/check-gate.js +26 -0
  25. package/proof/scripts/install-time.js +16 -0
  26. package/proof/scripts/lib.js +73 -0
  27. package/proof/scripts/measure.js +15 -0
  28. package/proof/scripts/missing-results.js +30 -0
  29. package/proof/scripts/record-gate.js +38 -0
  30. package/proof/scripts/render.js +18 -0
  31. package/proof/scripts/runner-overhead.js +21 -0
  32. package/src/README.md +9 -3
  33. package/src/activation-ownership.js +19 -0
  34. package/src/apply-companions.js +104 -0
  35. package/src/apply-snippets.js +60 -28
  36. package/src/aunx.js +262 -0
  37. package/src/catalog.js +253 -117
  38. package/src/install.js +478 -209
  39. package/src/plugin.js +13 -4
  40. package/src/postinstall.js +57 -0
  41. package/src/roles.js +184 -0
  42. package/src/uninstall.js +125 -8
  43. package/templates/README.md +19 -2
  44. package/templates/advanced/README.md +2 -2
  45. package/templates/advanced/vm/PRIVACY_GATES.md +17 -19
  46. package/templates/advanced/vm/README.md +25 -20
  47. package/templates/advanced/vm/box-CLAUDE.md +19 -18
  48. package/templates/advanced/vm/jobs/README.md +3 -1
  49. package/templates/advanced/vm/jobs/weekly-audit.service +3 -0
  50. package/templates/advanced/vm/jobs/weekly-audit.sh +2 -2
  51. package/templates/advanced/vm/setup-vm.sh +49 -2
  52. package/templates/agents/README.md +2 -2
  53. package/templates/agents/agy/README.md +20 -3
  54. package/templates/agents/agy/builder.md +11 -7
  55. package/templates/agents/agy/bulk-worker.md +9 -7
  56. package/templates/agents/agy/code-reviewer.md +13 -7
  57. package/templates/agents/agy/deep-planner.md +10 -7
  58. package/templates/agents/agy/done-verifier.md +13 -22
  59. package/templates/agents/agy/finding-verifier.md +14 -22
  60. package/templates/agents/agy/live-researcher.md +10 -7
  61. package/templates/agents/agy/reader.md +10 -12
  62. package/templates/agents/claude-code/README.md +18 -14
  63. package/templates/agents/claude-code/builder.md +10 -15
  64. package/templates/agents/claude-code/bulk-worker.md +8 -10
  65. package/templates/agents/claude-code/code-reviewer.md +11 -17
  66. package/templates/agents/claude-code/deep-planner.md +9 -11
  67. package/templates/agents/claude-code/done-verifier.md +12 -33
  68. package/templates/agents/claude-code/finding-verifier.md +13 -39
  69. package/templates/agents/claude-code/live-researcher.md +9 -11
  70. package/templates/agents/claude-code/reader.md +9 -18
  71. package/templates/agents/snippets/chat.md +9 -10
  72. package/templates/agents/snippets/claude-code.md +17 -18
  73. package/templates/agents/snippets/generic.md +9 -11
  74. package/templates/agents/snippets/route-gate.mjs +2 -2
  75. package/templates/agents/snippets/route-metrics.mjs +1 -1
  76. package/templates/agents/snippets/subagent-context.mjs +4 -4
  77. package/templates/beginner/ORCHESTRATOR.md +31 -36
  78. package/templates/beginner/README.md +1 -1
  79. package/templates/common/ACCEPTANCE_CHECKS.json +12 -0
  80. package/templates/common/CONTEXT.md +37 -0
  81. package/templates/common/DECISIONS.md +11 -0
  82. package/templates/common/README.md +24 -11
  83. package/templates/common/TASK_BRIEF.md +84 -0
  84. package/templates/common/protocols/README.md +14 -11
  85. package/templates/common/protocols/acceptance-checks.md +14 -0
  86. package/templates/common/protocols/build-protocol.md +91 -106
  87. package/templates/common/protocols/context-file.md +10 -0
  88. package/templates/common/protocols/decision-log.md +9 -0
  89. package/templates/common/protocols/deep-research.md +20 -34
  90. package/templates/common/protocols/docs-then-prove.md +13 -18
  91. package/templates/common/protocols/gap-analysis.md +15 -21
  92. package/templates/common/protocols/memory-and-record.md +21 -20
  93. package/templates/common/protocols/numbers-and-logic.md +20 -26
  94. package/templates/common/protocols/propagate.md +18 -27
  95. package/templates/intermediate/CLI-RUN.md +83 -113
  96. package/templates/intermediate/DELEGATION_MATRIX.md +9 -3
  97. package/templates/intermediate/README.md +3 -3
  98. package/templates/intermediate/RESEARCH_TRIAGE.md +23 -15
  99. package/templates/intermediate/ROUTING.md +54 -51
  100. package/templates/intermediate/TIERS.md +37 -76
  101. package/templates/tools/README.md +1 -1
  102. package/templates/tools/obsidian-tc/OBSIDIAN-TC.md +1 -1
  103. package/docs/audit-brief.md +0 -148
  104. package/scripts/README.md +0 -7
  105. package/scripts/gen-catalog.js +0 -81
  106. package/scripts/gen-plugin.js +0 -16
  107. package/scripts/record-demo.sh +0 -45
  108. package/templates/common/TASK_BUNDLE.md +0 -56
@@ -1,34 +1,25 @@
1
- # Propagate: change completeness
1
+ # Propagate: complete a shared change
2
2
 
3
- **A rename is a refactor, not a single-file edit.** Any change to a name, term, path, slug, schema field, routing rule or shared convention reaches everything that uses it, and the goal is zero silent strays.
3
+ When changing a name, term, path, slug, schema field, routing rule or shared convention, treat every consumer as part of the change. Use scoped search or a cheap worker to collect references before asking a planning model to decide anything.
4
4
 
5
- This is retrieval work. It stays with the orchestrator (or a cheap worker for the grep sweep). It never goes to the deep tier: a judgment model re-deriving a file list is the most expensive routing mistake there is.
5
+ ## Map consumers before editing
6
6
 
7
- ## 1. Map everything it touches (before editing anything)
7
+ - Search code, configuration, hooks, scheduled jobs and CI within the authorized repositories.
8
+ - Search project instructions and any authorized record store for the old term and link forms.
9
+ - Check documented external consumers such as webhooks, dashboards or generated outputs.
10
+ - When obsidian-tc is selected, use backlinks and text search; when absent, use file search and the project's own link checker.
11
+ - Scope recursive searches away from `.git`, `node_modules` and generated build output unless that output is explicitly part of verification.
8
12
 
9
- - **Docs and notes:** backlinks to the thing being renamed; literal search for the old term and its link forms. With obsidian-tc: `get_backlinks`, `search_text`, then `find_unresolved_links` after the change (`protocols/memory-and-record.md`).
10
- - **Memory / instructions:** grep every instructions file your agents read (`CLAUDE.md`, `AGENTS.md`, `GEMINI.md`, `QWEN.md`, custom instructions) and any memory store.
11
- - **Code / config:** grep the repos, settings files, hooks, scheduled jobs, CI, and the files this installer wrote.
12
- - **Other people's surfaces:** anything that consumes the old name from outside (webhooks, dashboards, bookmarks).
13
+ ## Update the affected surfaces
13
14
 
14
- ```bash
15
- grep -rniE "<old-term>" --exclude-dir=.git --exclude-dir=node_modules --exclude-dir=dist .
16
- ```
15
+ - Turn each hit into a checklist item with its owner and needed permission.
16
+ - Use a supported refactor or link-rewrite tool when one exists; otherwise apply scoped edits.
17
+ - Update indexes and documentation in the same change.
18
+ - When a consumer is outside the granted scope, return the exact required handoff and keep it visible in coverage.
17
19
 
18
- ## 2. Build the checklist
20
+ ## Verify old and new references
19
21
 
20
- Every hit becomes a line, grouped by surface. Nothing closes until the list is empty. Mark which items are approval-gated (production deploys, database migrations, anything public).
21
-
22
- ## 3. Execute on all fronts
23
-
24
- Change every item. Use the tool's own governed rename where one exists (a link rewriter, an IDE refactor) over hand edits. If an index or README lists the renamed thing, that index is part of the change, not a follow-up.
25
-
26
- ## 4. Verify: the loud negative (non-negotiable)
27
-
28
- Re-grep the OLD identifier across every surface. **Expect zero**, except historical records you name explicitly. Then check for dangling references the rename created (unresolved links, 404s, failing imports).
29
-
30
- Paste the final grep output as proof. If any stray survives, it is not done.
31
-
32
- ## Why the last step is the one that matters
33
-
34
- Steps 1 to 3 find what you thought of. Step 4 finds what you did not. A rename that "looks complete" and leaves one stray is worse than an unstarted one, because the stray is now believed.
22
+ - Search the old identifier again in every declared scope.
23
+ - Require zero active references, except named compatibility paths or historical records.
24
+ - Verify the new links, imports and consumer behavior with the relevant checker.
25
+ - Return the search command, exit status and any intentional retained matches as evidence.
@@ -1,165 +1,135 @@
1
- # CLI-RUN.md: exit 0 means a structurally accepted, non-empty response
1
+ # Lane runner (`aunx cli-run`)
2
2
 
3
- `bin/cli-run.mjs` is one entrypoint for the agent CLI lanes. It builds the right invocation per lane, reads that lane's native terminal event, and exits non-zero unless a structurally accepted, non-empty response came back.
3
+ When assigning work to a CLI lane, use `aunx cli-run` or the installed `bin/cli-run.mjs`. The runner builds the vendor invocation, reads its terminal event and returns an exit code for the response and any explicit output contract.
4
4
 
5
- ## The guarantee, exactly
5
+ Enabled lanes (edit `bin/lanes.json`): {{CLI_RUN_LANES}}.
6
6
 
7
- Exit 0 means: the lane's native terminal event says it finished, the response is non-empty, and the lane-specific error checks passed. **It does not mean the task was done.** A refusal that parses cleanly is exit 0. When your task has a real contract, state it:
7
+ ## Call a lane with a task brief
8
8
 
9
9
  ```bash
10
- node bin/cli-run.mjs codex "Write the report to out/report.md" --expect-file out/report.md # must exist, be non-empty, and be written during this run
11
- node bin/cli-run.mjs grok "Return the table as JSON" --expect-json # the response must parse as JSON
10
+ aunx cli-run {{EXAMPLE_LANE}} --brief TASK_BRIEF.md --timeout 900
11
+ node bin/cli-run.mjs {{EXAMPLE_LANE}} --brief TASK_BRIEF.md --timeout 900
12
12
  ```
13
13
 
14
- An unmet contract is exit 10 with reason `contract_unmet`. `--expect-file` snapshots the target before the lane starts (existence, size, mtime, content hash) and afterwards requires a non-empty regular file that is new or changed: a different hash, or a later mtime. A file that existed before and was not touched fails, however recent it is; a rewrite with identical bytes and an unchanged mtime also fails, because nothing distinguishes it from no write. Timestamps and hashes are evidence of change, not proof of authorship: if another process could write the same path during the run, use a per-attempt path. Text-only callers need nothing new: without a contract flag the behaviour is the structural check above.
14
+ When `aunx` runs, it uses the packaged runner. Use `--dir <rules-directory>` to select an installed project runner and its lane configuration explicitly. An absent project runner falls back to the package runner.
15
15
 
16
- Enabled lanes (edit `bin/lanes.json`): {{CLI_RUN_LANES}}
16
+ When requesting a particular route, inspect the lane's current roster and select the model and effort for the job:
17
17
 
18
18
  ```bash
19
- node bin/cli-run.mjs <grok|codex|agy|hermes|qwen> "<prompt>" [--brief FILE] [--timeout SECS] [--quiet]
20
- node bin/cli-run.mjs codex --audit "<prompt>" # read-only sandbox, the audit shape
21
- node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
22
- node bin/cli-run.mjs codex "<prompt>" --effort auto # medium or high, bounded heuristic
23
- node bin/cli-run.mjs qwen [--safe-mode] "<prompt>" # qwen-only flag
19
+ aunx cli-run {{EXAMPLE_LANE}} "<prompt>" --model '<model-id>'{{EXAMPLE_EFFORT_FLAGS}}
20
+ node bin/cli-run.mjs {{EXAMPLE_LANE}} "<prompt>" --model '<model-id>'{{EXAMPLE_EFFORT_FLAGS}}
24
21
  ```
25
22
 
26
- Put it on your PATH if you like: `ln -s "$PWD/bin/cli-run.mjs" ~/.local/bin/cli-run`.
23
+ {{EXAMPLE_AUDIT_BLOCK}}
27
24
 
28
- ## First run: `--doctor`
25
+ {{QWEN_SAFE_MODE_NOTE}}`--quiet` suppresses human-readable status lines. `--timeout SECS` bounds each call.
26
+
27
+ ## Check availability before work depends on it
29
28
 
30
29
  ```bash
31
- node bin/cli-run.mjs --doctor # which lanes are enabled and which binaries are on PATH; runs nothing
32
- node bin/cli-run.mjs --doctor --run # also sends each enabled lane one tiny prompt and judges the reply (uses a little quota)
30
+ aunx cli-run --doctor
31
+ node bin/cli-run.mjs --doctor
32
+ aunx cli-run --doctor --run
33
+ node bin/cli-run.mjs --doctor --run
33
34
  ```
34
35
 
35
- With zero enabled lanes, doctor reports inactive and exits 13. Use level 1 or select a supported CLI. Binary presence does not verify authentication or that your primary loaded its instructions.
36
-
37
- A lane that is enabled but not on PATH, or that answers with no deliverable, shows up here before it shows up mid-task.
38
-
39
- ## Why it exists
40
-
41
- Every agent CLI can exit 0 having produced nothing. The symptom (confident preamble, exit 0, no deliverable) is indistinguishable from a model failure, so it gets blamed on the model. Four wrong diagnoses in one week came from exactly that.
36
+ Use `--doctor` to inspect enabled lanes, binaries and requested defaults. Use `--doctor --run` for a small live prompt on each enabled lane, consuming vendor quota. When no lanes are enabled, the doctor exits 13; choose a supported CLI or use level 1. Verify authentication and loaded project instructions separately from binary presence.
42
37
 
43
- Byte count is not a deliverable check either: a run can emit hundreds of kilobytes and contain no conclusion.
38
+ ## Verify the result contract
44
39
 
45
- ## The success signal per lane
46
-
47
- | Lane | Invocation built | Success = |
48
- |---|---|---|
49
- | grok | `--output-format json -p` | `stopReason == "end_turn"` and non-empty `text` |
50
- | codex | `exec --json --color never --skip-git-repo-check -o FILE` | terminal `{"type":"turn.completed"}` and non-empty FILE |
51
- | agy | `--print-timeout Nm --output-format stream-json -p` | terminal `{"event":"result"}`, `status == "SUCCESS"`, non-empty `response` |
52
- | hermes | `-z … --usage-file FILE` | its exit code is already honest: 0 response · 1 none · 2 bad args |
53
- | qwen | `-o json [-m ID] [--safe-mode] -p` | terminal `{"type":"result"}`, `subtype == "success"`, `is_error` false, non-empty `result` not starting with `[API Error:`, and every `stats.models.*.api.totalErrors == 0` |
40
+ When exit 0 returns, the native terminal event indicates completion, the response is non-empty and lane-specific error checks passed. Verify the task's actual acceptance checks separately: a structurally valid refusal can still satisfy that response shape.
54
41
 
55
- qwen is the lane whose own success flags lie: an upstream 400 comes back as exit 0, `subtype: success`, `is_error: false`, with the error text inside `result`. The two extra checks are the honest ones. Absent telemetry is refused, not read as zero.
42
+ When the task requires a file or structured output, add an explicit contract:
56
43
 
57
- ## Exit codes
44
+ ```bash
45
+ aunx cli-run {{EXAMPLE_LANE}} "Write out/report.md" --expect-file out/report.md
46
+ node bin/cli-run.mjs {{EXAMPLE_LANE}} "Write out/report.md" --expect-file out/report.md
47
+ aunx cli-run {{EXAMPLE_LANE}} "Return the table as JSON" --expect-json
48
+ node bin/cli-run.mjs {{EXAMPLE_LANE}} "Return the table as JSON" --expect-json
49
+ ```
58
50
 
59
- | Code | Class | Meaning | What to do |
60
- |---|---|---|---|
61
- | 0 | `ok` | structurally accepted non-empty response, every `--expect-*` contract met | use it; if `refused=N` is above 0, read the problem line |
62
- | 10 | `empty` | ran and delivered nothing, or a contract was unmet | rerun once, or use another lane |
63
- | 11 | `no_output` | no output at all | rerun once; check the lane runs on its own |
64
- | 12 | `timeout` | timed out; the lane and every descendant in its process group were killed | raise `--timeout` or split the brief |
65
- | 13 | `unavailable` | binary missing, disabled in `lanes.json`, or `lanes.json` malformed | install or enable the lane |
66
- | 14 | `auth` | the lane's own error says a credential is missing or it is not logged in | set the credential; retrying cannot help |
67
- | 15 | `quota` | the lane's own error says usage limit, credits or rate limit | switch lanes or wait for the reset |
68
- | 16 | `rejected` | the upstream rejected the request: an unknown model id, a malformed request | fix the id, flag or request it names |
69
- | 17 | `refused` | no deliverable, and the lane reports tool calls a hook or deny rule blocked | adjust the rule, or give the lane the tool |
70
- | 18 | `cut_short` | no trustworthy finish: a missing or non-success terminal event, a lane killed by a signal, output past the 16 MiB buffer, or a nonzero vendor exit nothing above explains (hermes' own exit 2, bad args or an empty response, stays `empty`) | rerun once, then read the lane's stderr |
71
- | 130 / 143 | `interrupted` | cli-run itself received SIGINT / SIGTERM; the lane's process group was killed first, then the temp dir removed | |
72
- | 2 | | usage error in cli-run itself | |
51
+ `--expect-file` records existence, size, mtime and content hash before the run. It then requires a non-empty regular file that is new or changed by hash or later mtime. Use a unique per-attempt path when another process might write the same file. `--expect-json` requires parseable JSON. An unmet contract exits 10 with reason `contract_unmet`.
73
52
 
74
- A nonzero vendor exit is never `ok`, even when parseable text came back; the vendor's own code is kept in the log as `cli_rc`, and the bounded head of its stderr is shown on your terminal.
53
+ ## Read each lane's completion signal
75
54
 
76
- ## Why a run failed: the class
55
+ The runner supports these lanes whether or not you selected them.
77
56
 
78
- Exit 10 used to cover causes that need opposite responses. A missing API key is not a model fault, and retrying a spent quota cannot help. So every run lands in exactly one class, and on the terminal every failure (and every `ok` with refused calls) gets two more lines:
57
+ | Lane | Invocation built | Accepted response |
58
+ |---|---|---|
59
+ | grok | `--output-format json -p` | `stopReason == "end_turn"` and non-empty `text` |
60
+ | codex | `exec --json --color never --skip-git-repo-check -o FILE` | terminal `turn.completed` event and non-empty output file |
61
+ | agy | `--print-timeout Nm --output-format stream-json -p` | terminal `result` event, `status == "SUCCESS"`, non-empty `response` |
62
+ | hermes | `-z` with a usage file | vendor exit 0 and a non-empty response |
63
+ | qwen | `-o json`, optional model and safe-mode flags, `-p` | successful terminal result, no error flag or API-error result, non-empty text and every model's `api.totalErrors == 0` |
79
64
 
80
- ```
81
- cli-run[qwen] exit_nonzero rc=14 class=auth refused=0 0.8s raw=212B route=lane default :: lane exited 1; subtype="error_during_execution": Missing API key ...
82
- cli-run problem: cli-run[qwen] auth: Missing API key ...
83
- cli-run fix: set the credential the message above names (its environment variable, or the lane's own login command), then rerun
84
- ```
65
+ When Qwen's error telemetry is absent, the runner refuses the response. These checks distinguish response structure from successful task execution.
85
66
 
86
- A calling agent relays both lines and asks before fixing anything.
67
+ ## Respond to the exit class
87
68
 
88
- - **Signals come from each lane's authoritative error fields only.** codex: its own `error`, `turn.failed` and error-item events plus stderr. agy: the terminal result's status and error plus stderr. qwen: the terminal event's full error text (never the clipped display line) plus stderr. hermes: stderr on a nonzero exit, never its stdout. Never the model's prose, so an answer that explains what "rate limit" means is not a quota failure. grok and hermes have no native auth signal, and none is invented.
89
- - **Precedence** when several are present: `auth`, `quota`, `rejected`, `refused`, `cut_short`, `empty`. A missing credential explains everything downstream of it.
90
- - **`refused`** counts tool calls a hook or deny rule blocked: qwen's `permission_denials`, agy's deny-rule `TOOL_ERROR` steps, one per codex router `Rejected(` line on stderr, and grok's session transcript (`~/.grok/sessions/<cwd>/<sessionId>/updates.jsonl`: hook runs with `blocked`, and "Denied by permission policy"). grok's `sessionId` comes from lane output, so it must match `^[A-Za-z0-9-]{8,64}$`, the resolved path must stay inside the sessions root, every counted line must name that session, and the read stops at 5 MiB. `null` means the lane gave no readable signal (hermes always), which is an unknown, never 0.
91
- - **Refused with a deliverable is still exit 0.** The deliverable exists; `refused=N` and the problem and fix lines say what was blocked.
92
- - **Redacted.** The status, problem and stderr lines pass through one redaction pass before they are printed: JSON credential keys (escaped quotes, unterminated values and JSON escaped inside a string included), `Authorization:` values of any scheme, bearer values, URL query credentials, and the `sk-`, `xai-`, `ghp_` and `AIza` key prefixes. Redaction runs before any clipping, so a long token is never cut into an unrecognisable fragment. Denial text is searched in at most the first 256 KiB of stdout and of stderr, with bounded patterns, so hostile output cannot stall the wrapper after the lane has exited. None of these lines is ever logged.
69
+ | Code | Class | Action |
70
+ |---|---|---|
71
+ | 0 | `ok` | Verify task acceptance; read the problem line if `refused` is positive |
72
+ | 10 | `empty` | Inspect the missing response or unmet output contract |
73
+ | 11 | `no_output` | Check the lane directly before another attempt |
74
+ | 12 | `timeout` | Diagnose progress, then adjust the bound or split the work |
75
+ | 13 | `unavailable` | Resolve binary presence, lane enablement or malformed configuration |
76
+ | 14 | `auth` | Use the vendor's login or configured credential path |
77
+ | 15 | `quota` | Select another authorized lane or wait for quota renewal |
78
+ | 16 | `rejected` | Correct the model, flag or request named by the vendor |
79
+ | 17 | `refused` | Inspect the denied tool call and route it within authorized permissions |
80
+ | 18 | `cut_short` | Inspect a missing terminal event, signal, output overrun or unexplained nonzero vendor exit |
81
+ | 130 / 143 | `interrupted` | The wrapper received SIGINT / SIGTERM; inspect cleanup before resuming |
82
+ | 2 | usage error | Correct the runner arguments |
93
83
 
94
- ## The route: which model, and how hard it thinks
84
+ When a run fails, read the problem and fix lines and relay the actionable cause. Resolve it within existing authorization; request approval only for a change outside that scope. A nonzero vendor exit remains a failure even if text was produced; `cli_rc` preserves that vendor exit in the local log.
95
85
 
96
- A lane you do not pin runs on **its own config file**, which this tool cannot see. That is the quiet failure this section exists for: a CLI configured months ago at `reasoning_effort = "low"` keeps auditing at low effort while your routing docs describe a second-opinion pass, and nothing anywhere says so.
86
+ The runner classifies authoritative vendor error fields, with precedence `auth`, `quota`, `rejected`, `refused`, `cut_short`, `empty`. It excludes ordinary model prose from those error signals. A readable result with refused tool calls can remain exit 0; read the reported refusal alongside the result. `refused: null` means the lane supplied no readable refusal signal.
97
87
 
98
- Pin it per call, or per lane:
88
+ ## Record the requested route
99
89
 
100
- ```bash
101
- node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high # this call only
102
- node bin/cli-run.mjs --doctor # prints what each lane is pinned to
103
- ```
90
+ When a command flag is present, it overrides `bin/lanes.json` defaults. When both are absent, the lane uses its own configuration. Inspect that configuration when the model or effort matters.
104
91
 
105
92
  ```json
106
- {
107
- "enabled": ["codex", "grok"],
108
- "defaults": { "codex": { "model": "gpt-6-astra", "effort": "high" } }
109
- }
93
+ {{EXAMPLE_LANES_JSON}}
110
94
  ```
111
95
 
112
- A flag beats `defaults`; `defaults` beats nothing. Each vendor spells these differently and `cli-run` translates:
96
+ Replace the placeholder with a current vendor model ID before using this example. Model and effort values are bounded to the supported safe character set.
97
+
98
+ The runner supports these lanes whether or not you selected them.
113
99
 
114
- | Lane | Model | Reasoning effort |
100
+ | Lane | Model flag | Effort flag |
115
101
  |---|---|---|
116
102
  | grok | `-m` | `--reasoning-effort` |
117
103
  | codex | `-m` | `-c model_reasoning_effort="LEVEL"` |
118
- | agy | `--model` | `--effort` (low, medium, high) |
119
- | hermes | `-m` | `--reasoning` (none, minimal, ...) |
120
- | qwen | `-m` | none: this lane has no reasoning flag |
121
-
122
- Three rules that keep this honest:
123
-
124
- - **A level `cli-run` does not recognise is not rejected here.** Levels are the vendor's, they change, and guessing the valid set would date this tool. An unknown level is refused by the lane and surfaces as a class (codex reports it as `rejected`, exit 16; a lane with no rejection signal as `cut_short`, exit 18), with the lane's own exit code in the log as `cli_rc` and its stderr on your terminal.
125
- - **`--effort` on qwen is a usage error, not a silent drop.** A flag that vanishes leaves you believing a route that never ran.
126
- - **`--effort auto` is bounded.** It uses prompt size, or an audit's changed-file evidence, and resolves only medium or high. An audit is always high. Auto is a heuristic, not a measurement: name `xhigh` explicitly for security-critical or irreversible work.
127
- - **Values are charset-bounded** (letters, digits, and `. _ : @ / + -`, no leading dash, 64 characters). A model id becomes an argv element and, on codex, part of a TOML value; bounding it is what stops either from being escaped.
128
-
129
- ## Permissions are a separate layer
130
-
131
- `cli-run` never injects permission flags. Each CLI carries its own config, so every caller gets the same behaviour. Use each vendor's deny-list as the base layer; allow-lists only hold if every binary is enumerable in advance.
132
-
133
- ## Log
134
-
135
- `~/.ai-orchestrator/cli-run.log.jsonl`, one line per run: lane, verdict, `class` (one of the classes above), rc, the lane's own exit code (`cli_rc`), signal, `refused` (an integer, or `null` when the lane gives no signal), seconds, raw bytes, deliverable bytes, a 12-hex sha256 prefix of the prompt and its length, the route (`model_requested`, `effort_requested`, and `model_source` / `effort_source`, each one of `flag`, `lanes.json` or `lane_default`), auto-sizing evidence (`effort_resolved`, `effort_basis`, `effort_scope`), and `reason`: one of a fixed set of codes (`ok`, `not_json`, `bad_stop_reason`, `empty_text`, `no_terminal_event`, `bad_status`, `api_error_in_result`, `total_errors`, `contract_unmet`, `exit_nonzero`, `timeout`, `killed`, `disabled`, `lanes_json_malformed`, ...). `effort_basis` is exactly one of `explicit`, `prompt_chars`, `audit_floor`, or `none`. Never the prompt text, never a provider-supplied value, never free text: a value the log does not recognise is written as `unknown`. The human-readable detail, which may quote the provider, goes to your terminal only (and nowhere with `--quiet`). "This lane is flaky" becomes a query instead of an argument, and so does "we route audits at high effort".
136
-
137
- The log records what was **requested**, on every record including a run refused before the lane started. It does not record an actual. Reporting is inconsistent: grok returns a `modelUsage` block naming a model, the other four lanes return nothing of the kind, so an `actual` field would be populated for one lane and empty for four. It would also be a provider-supplied string, and this log holds fixed codes and bounded caller-supplied values only. `model_source: "lane_default"` is the honest way to say this run inherited something invisible from here.
138
-
139
- ## The prompt travels in argv
140
-
141
- That is each vendor's documented headless shape (`-p`, `exec`). Two consequences: argv is visible to other processes on the machine, so a prompt is never the place for a key; and argv is bounded by the OS (`ARG_MAX`), so a very large brief should be referenced by path inside the prompt rather than pasted whole.
104
+ | agy | `--model` | `--effort` |
105
+ | hermes | `-m` | `--reasoning` |
106
+ | qwen | `-m` | Unsupported; an effort request is a usage error |
142
107
 
143
- ## lanes.json refuses by default
108
+ When using `--effort auto`, treat its medium/high selection as a bounded heuristic; an audit has a high floor. Use explicit high for builds and xhigh where supported for security-critical or irreversible work. The vendor validates its own effort names and reports unsupported values through the failure class.
144
109
 
145
- Absent: every lane enabled, nothing pinned. Present but malformed or unreadable: every lane refused (exit 13) until it is fixed. A half-written config never re-enables a lane the installer disabled. `defaults` is optional and held to the same standard: a malformed entry, an unknown lane, an unknown key, a value outside the charset, or an effort pinned on a lane that has no reasoning flag all make the whole file refuse by default rather than being skipped quietly.
110
+ ## Preserve permissions and secrets
146
111
 
147
- ## A killed lane is not a deliverable
112
+ The runner never adds permission flags. Keep each vendor's permissions in its own configuration and route work within the task's granted scope. A denied write becomes a handoff to an authorized writer.
148
113
 
149
- A lane that dies by signal has no honest exit status. Whatever it printed first is discarded; the run reports `killed`, class `cut_short`, exit 18.
114
+ Prompts travel in argv, which other processes may inspect. Never put secrets in a prompt. For large task inputs, give the worker a brief with authorized source paths instead of exceeding the operating system's argument-size limit.
150
115
 
151
- ## Interrupting cli-run kills the lane too
116
+ Terminal status and error details are redacted before clipping. Local logs exclude prompt text and provider free text. `--quiet` also suppresses the human-readable details.
152
117
 
153
- Ctrl-C or a `kill` on the wrapper kills the lane's whole process group before the wrapper exits (130 for SIGINT, 143 for SIGTERM). A second signal during cleanup kills again and exits at once. Handlers are installed per run and removed when it finishes, so `--doctor --run` does not accumulate them. Uncatchable SIGKILL to the wrapper leaves the lane running; that is the operating system, not a promise this tool can make. Under systemd, `KillMode=control-group` covers that case.
118
+ ## Inspect the local log
154
119
 
155
- ## Output is decoded as a UTF-8 stream
120
+ Read `~/.ai-orchestrator/cli-run.log.jsonl` for one record per run: lane, verdict, class, return codes, signal, refusal count, timing, byte counts, prompt hash and length, requested model and effort, source of those requests, auto-effort evidence and a fixed reason code.
156
121
 
157
- Vendor output is decoded with a streaming decoder, so a multibyte character split across two chunks is preserved byte for byte. The 16 MiB cap and the `raw_bytes` field count bytes, not characters.
122
+ Use `model_source: "lane_default"` to identify an inherited route. These fields record requested settings; verifying the vendor's actual model requires vendor evidence. Keep task success grounded in the acceptance checks.
158
123
 
159
- ## Timeouts kill the whole process group
124
+ ## Configuration and process boundaries
160
125
 
161
- The lane is started detached, as the leader of its own process group. On timeout, or when output overruns the buffer, the group is killed, so a tool the agent shelled out to cannot keep writing after the wrapper reported 12. A child that calls `setsid()` itself escapes this boundary; nothing user-space can promise more without a cgroup, which is what the level 3 systemd unit adds.
126
+ - When `lanes.json` is absent, all supported lanes are enabled with inherited defaults. When it is malformed or unreadable, the runner refuses every lane with exit 13 until corrected.
127
+ - When a lane dies by signal, discard its partial response and handle `cut_short` exit 18.
128
+ - On POSIX, timeout, output overrun and catchable interruption kill the lane's process group. A child that creates its own session can escape that boundary; use a service-level process boundary where needed.
129
+ - On Windows, lanes use a direct executable or a resolved Node shim, never `cmd.exe`. Windows termination behavior differs from POSIX signal cleanup.
130
+ - When the wrapper receives uncatchable SIGKILL, use a supervisor such as systemd with `KillMode=control-group` to clean up its process tree.
131
+ - Vendor output uses streaming UTF-8 decoding and a 16 MiB cap counted in bytes.
162
132
 
163
- ## Lane choice is not automated
133
+ ## Choose the lane
164
134
 
165
- `cli-run` runs the lane it is given. Which lane fits the job is `ROUTING.md` and `DELEGATION_MATRIX.md`, or a question to the human.
135
+ When selecting a lane, apply `ROUTING.md` and `DELEGATION_MATRIX.md`, optionally starting with `aunx route "<task>"`. The runner executes the lane named by the caller. Its response checks work independently of optional companion software.
@@ -1,7 +1,13 @@
1
- # DELEGATION_MATRIX.md: task → lane → pick
1
+ # DELEGATION_MATRIX.md: choose a lane for the task
2
2
 
3
3
  Generated {{DATE}} from the AIs you said you have: `{{AI_IDS}}`.
4
4
 
5
+ {{STACK_TABLE}}
6
+
7
+ {{STACK_FALLBACK_NOTE}}
8
+
9
+ {{STACK_GAPS}}
10
+
5
11
  ## Your lanes
6
12
 
7
13
  {{LANES_TABLE}}
@@ -10,7 +16,7 @@ Generated {{DATE}} from the AIs you said you have: `{{AI_IDS}}`.
10
16
 
11
17
  {{PLAN_GUIDANCE}}
12
18
 
13
- ## Task → lane
19
+ ## Match the task to an available lane
14
20
 
15
21
  {{TASK_LANES_TABLE}}
16
22
 
@@ -24,4 +30,4 @@ Generated {{DATE}} from the AIs you said you have: `{{AI_IDS}}`.
24
30
 
25
31
  ## Privacy gate
26
32
 
27
- No private notes, client data, or personal records go to a metered third-party bulk lane or a fan-out lane. Name the barred lanes explicitly in your own rules; an unnamed bar is not enforced.
33
+ Before dispatching protected data, check the project's named allowed and barred lanes. Never send private notes, client data or personal records to an unapproved third-party bulk or fan-out lane. When the required permission or tool is absent, keep the data local and return the needed decision.
@@ -5,9 +5,9 @@ Written at level 2 and above, on top of `common/` and `beginner/`.
5
5
  | File | What it adds |
6
6
  |---|---|
7
7
  | `ROUTING.md` | the multi-lane decision tree; supersedes `ORCHESTRATOR.md` when present |
8
- | `TIERS.md` | capability tiers, the three cost levers, the escalation rule, slotting rules |
8
+ | `TIERS.md` | capability tiers, the three cost levers, job fit, model availability and effort selection |
9
9
  | `DELEGATION_MATRIX.md` | task → lane → pick, generated from the user's selection |
10
- | `RESEARCH_TRIAGE.md` | three engines in parallel, one triager |
11
- | `CLI-RUN.md` | how `bin/cli-run.mjs` judges each lane |
10
+ | `RESEARCH_TRIAGE.md` | selected engines in parallel, one source-checking writer |
11
+ | `CLI-RUN.md` | how `aunx cli-run` and `bin/cli-run.mjs` verify each response |
12
12
 
13
13
  `bin/cli-run.mjs` and `bin/lanes.json` are written by the installer from `bin/cli-run.mjs` in this repo and the user's selection; they are not templates.
@@ -1,30 +1,38 @@
1
- # RESEARCH_TRIAGE.md: engines in parallel, one triager
1
+ # RESEARCH_TRIAGE.md: parallel research with source verification
2
2
 
3
- {{RESEARCH_SELECTION_ADVICE}} Then triage against primary sources you open yourself.
3
+ When a question needs outside evidence, {{RESEARCH_SELECTION_ADVICE}} Then open the primary sources that support decisions.
4
4
 
5
- Your `cli-run` lanes: {{CLI_RUN_LANES}} ({{RESEARCH_ENGINES}} research engine(s) below). Everything in this file was rendered from that selection; a lane that is not listed is not one you have.
5
+ Use the selected `cli-run` lanes: {{CLI_RUN_LANES}} ({{RESEARCH_ENGINES}} research engine(s) below). When a needed lane is absent, use an available authorized research tool and name the resulting coverage limit.
6
6
 
7
- ## Roles
7
+ ## Assign the research roles
8
8
 
9
9
  | Role | Typical lane | Job |
10
10
  |---|---|---|
11
- {{RESEARCH_ROLES}} opens primary sources, marks every claim, writes the artifact |
11
+ {{RESEARCH_ROLES}}
12
12
 
13
- Run each engine as one `cli-run` call with a task bundle in `--brief`. A run that produced nothing exits 10 and is a missing engine, not an empty finding.
13
+ Give each engine a task brief through `--brief`, with the same context file, bounded questions, source standard and stopping condition. When a run returns no usable result, record that engine as unavailable and continue independent source checks.
14
14
 
15
- ## One run
15
+ ## Run the selected engines
16
16
 
17
17
  ```bash
18
- BRIEF=research/brief.md # purpose, sub-questions, source standard, report contract, exit parameters
18
+ BRIEF=research/brief.md
19
19
  {{RESEARCH_RUN}}
20
20
  ```
21
21
 
22
- Then the orchestrator reads the available outputs, opens every primary source that carries a decision, and writes one dated brief with marks: **CONFIRMED** (two engines + primary source) · **DISAGREEMENT** (both readings kept) · **REPORTED** (someone's own post, quoted not trusted) · **UNVERIFIED**.
22
+ When using the installed script directly, `aunx cli-run` accepts the same runner arguments as `node bin/cli-run.mjs`.
23
23
 
24
- ## Triage discipline
24
+ When outputs arrive, have one writer inspect the primary sources carrying each decision and produce a dated synthesis:
25
25
 
26
- - Plant one deliberately wrong figure in one brief. An engine that does not correct it has confirmations worth less than they look.
27
- - Expect one engine to return confident unsourced numerics and claim full coverage. Downgrade to hypothesis. Weight the engines that report their own gaps.
28
- - Agreement is weak evidence. Disagreement is the signal.
29
- - Only the orchestrator writes the durable record. Every other engine proposes.
30
- - Count dispositions, not briefs.
26
+ - **CONFIRMED:** verified against the primary source, with independent corroboration where the question needs it.
27
+ - **DISAGREEMENT:** preserve conflicting readings and identify the evidence that would resolve them.
28
+ - **REPORTED:** attribute a source's statement to that source.
29
+ - **UNVERIFIED:** name the missing evidence or access.
30
+
31
+ ## Verify claims and coverage
32
+
33
+ - When engines agree, verify the shared premise against source.
34
+ - When engines disagree, preserve both claims until evidence resolves them.
35
+ - When a report gives a number, open its source or recompute it with an available tool.
36
+ - When testing a research method, use a labelled false-premise fixture and confirm the method rejects it before relying on it.
37
+ - When recording results, use one writer and return claim dispositions with citations.
38
+ - When Context7 or another companion is absent, use official documentation, source and the project's own runtime checks.
@@ -1,8 +1,16 @@
1
- # ROUTING.md: the multi-lane decision tree
1
+ # ROUTING.md: model router for your available tools
2
2
 
3
- Primary agent (the orchestrator): **{{PRIMARY_NAME}}**. It routes, maps, builds, verifies and records. Every other AI is a lane it calls.
3
+ Main agent: **{{PRIMARY_NAME}}**. A **lane** is an AI tool or model the main agent can hand work to. A **tier** describes a model's capability and cost: planning model, working model or cheap model.
4
4
 
5
- Your lanes:
5
+ When a task arrives, choose its route from the live tools and these rules. `aunx route "<task>"` prints a deterministic keyword suggestion; verify that suggestion against the task's scope, required tools and stakes before dispatching.
6
+
7
+ {{STACK_TABLE}}
8
+
9
+ {{STACK_FALLBACK_NOTE}}
10
+
11
+ {{STACK_GAPS}}
12
+
13
+ ## Your lanes
6
14
 
7
15
  {{LANES_TABLE}}
8
16
 
@@ -10,76 +18,71 @@ Your lanes:
10
18
 
11
19
  {{PLAN_GUIDANCE}}
12
20
 
13
- Two kinds of lane. **Lane A** = subscription CLIs: $0 marginal, already paid for, used for interactive and agentic work. **Lane B** = metered APIs: per token, used for programmatic bulk where a subscription CLI cannot serve. **Local** = stays on the machine; a privacy lane, never a cost lane.
21
+ When choosing a billing route, distinguish **subscription lanes** (usage covered by a subscription, subject to its limits) from **pay-per-token lanes** (metered API usage). When data must stay on the machine, choose a verified **local lane** and check its network behavior.
14
22
 
15
- Rule of thumb: never spend a frontier token on a task a cheap tier finishes correctly. Escalate on signal (low confidence, explicit complexity, a failed verification), not by default. And an external lane must earn the hop with a real strength; when in doubt, stay in-house.
23
+ When a cheaper eligible route can complete and verify the task, select it. When verification fails, complexity emerges or tool reach is insufficient, escalate with a named reason.
16
24
 
17
25
  ## Decision tree (first match wins)
18
26
 
19
- 0. **Is there a cheaper or better external lane for this?** Check `DELEGATION_MATRIX.md`. Your enabled lanes, every one called through `bin/cli-run.mjs`:
27
+ 0. **An external lane offers useful reach or capacity:** check `DELEGATION_MATRIX.md` and call the selected CLI through `aunx cli-run` or the installed `bin/cli-run.mjs`.
20
28
  {{LANE_STEP0}}
21
- 1. **Bulk and mechanical?** → fast tier{{BULK_LANE}}.{{FAN_OUT_ADVICE}}
22
- 1a. **Reading or digesting many files or notes, not writing?** → reader. Different from a bulk pass: reader reports, it does not classify, tag or transform.
23
- 2. **Needs live data?** → {{LIVE_LANE}} standard tier with web tools.
24
- 3. **Reviewing without changing?** → standard tier read-only. Security-critical → {{ATTACK_LANE}}.
25
- 3a. **Holding findings from a review or a scanner?** → finding-verifier before any of them cause a repair. A finding is a claim, not a fact.
26
- 3b. **Checking a tracker item or task against its stated done-signal?** → done-verifier. It probes the named artifact and returns MET, NOT_MET or UNVERIFIABLE; it never closes anything itself.
27
- 4. **Ambiguous, strategic, expensive to get wrong?** → deep tier (deep-planner). Then hand the plan down.
29
+ 1. **Bulk or mechanical work:** classify, tag, extract, rename or reformat -> cheap model tier / {{BULK_ROLE}}{{BULK_LANE}}.{{FAN_OUT_ADVICE}}
30
+ 1a. **Read or digest many files:** -> {{READER_ROLE}}; return facts, quotes or an index within the brief's scope.
31
+ 2. **Current data is required:** -> {{LIVE_LANE}} working model tier with live tools.
32
+ 3. **Review code without changing it:** -> {{REVIEW_ROLE}}, working model tier. For security-critical scope -> {{ATTACK_LANE}} with appropriate effort.
33
+ 3a. **A reviewer or scanner has returned findings:** -> {{FINDING_ROLE}}; reproduce each claim before repair.
34
+ 3b. **Check a task's definition of done:** -> {{DONE_ROLE}}; probe the named artifact and return MET, NOT_MET or UNVERIFIABLE.
35
+ 4. **Ambiguity, architecture or an unknown cause:** -> planning model tier / {{PLANNER_ROLE}}. Return a concrete plan for execution.
28
36
  {{DECISION_RULE5}}
29
37
 
30
38
  {{WHO_BUILDS}}
31
39
 
32
- ## The Build Protocol, with lanes bound
40
+ ## Build process and ownership
33
41
 
34
- | Stage | Binding |
35
- |---|---|
36
- | 0 Route | live probe for access; `cli-run` lanes are $0 and uncapped |
37
- | 1 Map | the orchestrator sweeps{{STAGE1_LANES}} |
38
- | 2 Judge | deep tier, on the finished map: one named weak spot and one gap in the request |
39
- | 3 Build | the orchestrator, against the installed dependency's source |
40
- | 4 Scan | secret + static + dependency scanners, diff-scoped, refuses by default |
41
- | 5 Challenge | security-shaped diff → {{ATTACK_LANE}}. Architecture-shaped → deep tier, build against plan. Never both |
42
- | 5a Verify findings | finding-verifier, a different model family where you have one: CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Only CONFIRMED earns a repair |
43
- | 5b Ship | rollback id recorded, explicit human yes |
44
- | 6 Verify | real test, negative test seen red, old identifier re-grepped to zero |
45
- | 7 Record | one end-to-end doc, tracker Done with evidence, plan doc deleted |
42
+ When a task builds or changes a system, run `protocols/build-protocol.md`.
46
43
 
47
- Caps: two deep-tier checkpoints per build. CLI lanes are $0 and do not count.
48
-
49
- ## Numbers and logic
50
-
51
- Every number, comparison, complexity or equivalence claim goes through a tool that computes (`protocols/numbers-and-logic.md`; companion: codecalc, {{CODECALC_STATUS}}).{{METERED_CITATION_NOTE}}
52
-
53
- ## Memory and record
44
+ | Step | Action |
45
+ |---|---|
46
+ | Frame and probe | Quote the ask, freeze acceptance checks and verify required availability |
47
+ | Ordering | Dependencies first, invalidators early, deterministic checks before judgment, irreversible actions last |
48
+ | Research | Vet sources, inspect current interfaces and spike risky assumptions |
49
+ | Context file | Map affected surfaces once; every brief reads the same context file |
50
+ | Assign | Choose each section's lane, model and effort by live capability and job fit |
51
+ | Build | Execute the approved scope; hand refused writes to an authorized writer |
52
+ | Split and merge | Keep the whole scope in each brief; merge sections and name conflicts |
53
+ | Audit | One pass on the merged artifact plus a companion consult asking scope versus ask |
54
+ | Ship | Record rollback, confirm authorization, replay checks and verify the change in use |
55
+ | Record | Return coverage and evidence, update documentation and name what watches it |
54
56
 
55
- One writer per run; every other lane proposes. Search before writing, index in the same pass (`protocols/memory-and-record.md`; companion, optional: obsidian-tc, {{OBSIDIAN_TC_STATUS}}).
57
+ When work runs in the background, check liveness and output growth every five minutes. Two checks without growth call for diagnosis and a report. When an audit finding is confirmed, assign its fix to a non-author and show the regression failing before the fix. Verify the fix; no second audit pass.
56
58
 
57
- ## Docs, then prove
59
+ ## Tools and fallbacks
58
60
 
59
- A lane's recall of a library's API is a lead, not a verdict, the same as its arithmetic. Pull current, version-specific docs before writing a call against anything you have not confirmed this session (`protocols/docs-then-prove.md`; companion, optional: Context7, {{CONTEXT7_STATUS}}). Then prove the doc was right by running it, the same tool that already owns numbers: codecalc, {{CODECALC_STATUS}}. Where the two disagree, the run wins.
61
+ - When reporting consequential arithmetic or code equivalence, use a computing tool (`protocols/numbers-and-logic.md`). codecalc: {{CODECALC_STATUS}}. When absent, use the local runtime, test suite or spreadsheet.{{METERED_CITATION_NOTE}}
62
+ - When writing durable records, search first, update the index and keep one writer (`protocols/memory-and-record.md`). obsidian-tc: {{OBSIDIAN_TC_STATUS}}. When absent, use project files, search and version control.
63
+ - When using a changing library or API, read current documentation and verify behavior (`protocols/docs-then-prove.md`). Context7: {{CONTEXT7_STATUS}}. When absent, read official docs or installed source; use the local runtime when codecalc is absent.
60
64
 
61
- ## Modifier rules
65
+ ## Choose effort and verify the route
62
66
 
63
67
  {{PLAN_BIG_LINE}}{{INLINE_THRESHOLD_NOTE}}
64
- - **Escalation:** never silently retry at the same tier. Escalate one tier or consult deep once, and say which. Two consults that do not unstick it → stop and tell the human.
65
- - **De-escalation:** a request that sounds deep but is a lookup routes down.
66
- - **Long context:** mechanical digestion → fast tier in chunks; judgment over a long input → standard tier.
67
- - **Token discipline on every delegation:** pass only the context the delegate needs, never the conversation.
68
- - **Effort per agent:** deep xhigh, review, verification and build high, live research medium, bulk low.
69
- - **Three inputs, not one:** role picks the agent, complexity moves the effort, stakes move the tier and who reads it. A one-line auth change is simple and high-stakes at once, and the stakes decide. See `TIERS.md`.
70
- - **Pin the route when it matters:** a lane with no `--model`/`--effort` and no `defaults` entry in `bin/lanes.json` runs on its own config, which may be nothing like what this file describes. `cli-run --doctor` prints what each lane is pinned to, and every run logs the value requested and where it came from.
68
+ - When a lookup is sufficient, use the cheap model tier; when the task needs judgment, choose the working or planning model tier by evidence.
69
+ - When an attempt fails, identify the failure class before retrying; change the route or resolve the cause explicitly.
70
+ - When delegating, pass the context file and task brief with scoped source references and acceptance checks.
71
+ - When role, complexity or stakes change, reassess model, effort and reviewer together. See `TIERS.md`.
72
+ - Before each build, inspect the lane's live roster and any configured pin. Flag stale pins and record the model and effort selected for this task.
73
+ - When the route matters, set `--model` and `--effort` or verify `bin/lanes.json` defaults. Run `aunx cli-run --doctor` to inspect requested defaults; an unpinned lane uses its own configuration.
71
74
 
72
75
  ## Example routings
73
76
 
74
77
  | Task | Route |
75
78
  |---|---|
76
- | "Design the architecture for X" | deep-planner |
77
- | "Review this service for bugs" | code-reviewer |
79
+ | Design the architecture for a service | {{PLANNER_ROLE}}, planning model tier |
80
+ | Review this service for bugs | {{REVIEW_ROLE}}, working model tier |
78
81
  {{ADD_ENDPOINT_ROW}}
79
- | "Why does this silently drop rows sometimes" | deep-planner (unknown cause), then build the fix directly |
80
- | "Summarize these 30 notes into one index" | bulk-worker |
81
- | "Read every note in this folder and pull out every mention of X" | reader |
82
- | "The audit returned 6 findings" | finding-verifier first; repair only what comes back CONFIRMED |
83
- | "Is issue #123 actually done" | done-verifier |
82
+ | Find why this silently drops rows | {{PLANNER_ROLE}}, then a scoped build |
83
+ | Summarize similar notes into one index | {{BULK_ROLE}}, cheap model tier |
84
+ | Read every file and extract mentions of a topic | {{READER_ROLE}}, cheap model tier |
85
+ | Verify the audit's findings | {{FINDING_ROLE}} before repairs |
86
+ | Check whether the stated definition of done holds | {{DONE_ROLE}} |
84
87
  {{LANE_EXAMPLES}}
85
88
  {{ROUTE_GATE_SECTION}}