model-orchestrator 0.1.34 → 1.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +31 -21
- package/CHANGELOG.md +51 -1
- package/README.md +127 -110
- package/bin/README.md +57 -6
- package/bin/aunx.js +7 -0
- package/bin/cli-run.mjs +21 -15
- package/bin/cli.js +376 -257
- package/docs/README.md +15 -18
- package/docs/catalog.md +228 -38
- package/docs/companions.md +28 -10
- package/docs/guarantees.md +21 -12
- package/docs/how-it-routes.md +49 -42
- package/docs/install.md +135 -33
- package/docs/part-1-beginner.md +37 -45
- package/docs/part-2-intermediate.md +34 -52
- package/docs/part-3-advanced.md +36 -26
- package/docs/security-review-history.md +38 -0
- package/llms.txt +24 -25
- package/package.json +16 -8
- package/proof/README.md +100 -0
- package/proof/gate-demo.cast +9 -0
- package/proof/gate-demo.gif +0 -0
- package/proof/results.json +198 -0
- package/proof/scripts/check-gate.js +26 -0
- package/proof/scripts/install-time.js +16 -0
- package/proof/scripts/lib.js +73 -0
- package/proof/scripts/measure.js +15 -0
- package/proof/scripts/missing-results.js +30 -0
- package/proof/scripts/record-gate.js +38 -0
- package/proof/scripts/render.js +18 -0
- package/proof/scripts/runner-overhead.js +21 -0
- package/src/README.md +9 -3
- package/src/activation-ownership.js +19 -0
- package/src/apply-companions.js +104 -0
- package/src/apply-snippets.js +60 -28
- package/src/aunx.js +262 -0
- package/src/catalog.js +253 -117
- package/src/install.js +478 -209
- package/src/plugin.js +13 -4
- package/src/postinstall.js +57 -0
- package/src/roles.js +184 -0
- package/src/uninstall.js +125 -8
- package/templates/README.md +19 -2
- package/templates/advanced/README.md +2 -2
- package/templates/advanced/vm/PRIVACY_GATES.md +17 -19
- package/templates/advanced/vm/README.md +25 -20
- package/templates/advanced/vm/box-CLAUDE.md +19 -18
- package/templates/advanced/vm/jobs/README.md +3 -1
- package/templates/advanced/vm/jobs/weekly-audit.service +3 -0
- package/templates/advanced/vm/jobs/weekly-audit.sh +2 -2
- package/templates/advanced/vm/setup-vm.sh +49 -2
- package/templates/agents/README.md +2 -2
- package/templates/agents/agy/README.md +20 -3
- package/templates/agents/agy/builder.md +11 -7
- package/templates/agents/agy/bulk-worker.md +9 -7
- package/templates/agents/agy/code-reviewer.md +13 -7
- package/templates/agents/agy/deep-planner.md +10 -7
- package/templates/agents/agy/done-verifier.md +13 -22
- package/templates/agents/agy/finding-verifier.md +14 -22
- package/templates/agents/agy/live-researcher.md +10 -7
- package/templates/agents/agy/reader.md +10 -12
- package/templates/agents/claude-code/README.md +18 -14
- package/templates/agents/claude-code/builder.md +10 -15
- package/templates/agents/claude-code/bulk-worker.md +8 -10
- package/templates/agents/claude-code/code-reviewer.md +11 -17
- package/templates/agents/claude-code/deep-planner.md +9 -11
- package/templates/agents/claude-code/done-verifier.md +12 -33
- package/templates/agents/claude-code/finding-verifier.md +13 -39
- package/templates/agents/claude-code/live-researcher.md +9 -11
- package/templates/agents/claude-code/reader.md +9 -18
- package/templates/agents/snippets/chat.md +9 -10
- package/templates/agents/snippets/claude-code.md +17 -18
- package/templates/agents/snippets/generic.md +9 -11
- package/templates/agents/snippets/route-gate.mjs +2 -2
- package/templates/agents/snippets/route-metrics.mjs +1 -1
- package/templates/agents/snippets/subagent-context.mjs +4 -4
- package/templates/beginner/ORCHESTRATOR.md +31 -36
- package/templates/beginner/README.md +1 -1
- package/templates/common/ACCEPTANCE_CHECKS.json +12 -0
- package/templates/common/CONTEXT.md +37 -0
- package/templates/common/DECISIONS.md +11 -0
- package/templates/common/README.md +24 -11
- package/templates/common/TASK_BRIEF.md +84 -0
- package/templates/common/protocols/README.md +14 -11
- package/templates/common/protocols/acceptance-checks.md +14 -0
- package/templates/common/protocols/build-protocol.md +91 -106
- package/templates/common/protocols/context-file.md +10 -0
- package/templates/common/protocols/decision-log.md +9 -0
- package/templates/common/protocols/deep-research.md +20 -34
- package/templates/common/protocols/docs-then-prove.md +13 -18
- package/templates/common/protocols/gap-analysis.md +15 -21
- package/templates/common/protocols/memory-and-record.md +21 -20
- package/templates/common/protocols/numbers-and-logic.md +20 -26
- package/templates/common/protocols/propagate.md +18 -27
- package/templates/intermediate/CLI-RUN.md +83 -113
- package/templates/intermediate/DELEGATION_MATRIX.md +9 -3
- package/templates/intermediate/README.md +3 -3
- package/templates/intermediate/RESEARCH_TRIAGE.md +23 -15
- package/templates/intermediate/ROUTING.md +54 -51
- package/templates/intermediate/TIERS.md +37 -76
- package/templates/tools/README.md +1 -1
- package/templates/tools/obsidian-tc/OBSIDIAN-TC.md +1 -1
- package/docs/audit-brief.md +0 -148
- package/scripts/README.md +0 -7
- package/scripts/gen-catalog.js +0 -81
- package/scripts/gen-plugin.js +0 -16
- package/scripts/record-demo.sh +0 -45
- package/templates/common/TASK_BUNDLE.md +0 -56
|
@@ -1,34 +1,25 @@
|
|
|
1
|
-
# Propagate: change
|
|
1
|
+
# Propagate: complete a shared change
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
When changing a name, term, path, slug, schema field, routing rule or shared convention, treat every consumer as part of the change. Use scoped search or a cheap worker to collect references before asking a planning model to decide anything.
|
|
4
4
|
|
|
5
|
-
|
|
5
|
+
## Map consumers before editing
|
|
6
6
|
|
|
7
|
-
|
|
7
|
+
- Search code, configuration, hooks, scheduled jobs and CI within the authorized repositories.
|
|
8
|
+
- Search project instructions and any authorized record store for the old term and link forms.
|
|
9
|
+
- Check documented external consumers such as webhooks, dashboards or generated outputs.
|
|
10
|
+
- When obsidian-tc is selected, use backlinks and text search; when absent, use file search and the project's own link checker.
|
|
11
|
+
- Scope recursive searches away from `.git`, `node_modules` and generated build output unless that output is explicitly part of verification.
|
|
8
12
|
|
|
9
|
-
|
|
10
|
-
- **Memory / instructions:** grep every instructions file your agents read (`CLAUDE.md`, `AGENTS.md`, `GEMINI.md`, `QWEN.md`, custom instructions) and any memory store.
|
|
11
|
-
- **Code / config:** grep the repos, settings files, hooks, scheduled jobs, CI, and the files this installer wrote.
|
|
12
|
-
- **Other people's surfaces:** anything that consumes the old name from outside (webhooks, dashboards, bookmarks).
|
|
13
|
+
## Update the affected surfaces
|
|
13
14
|
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
15
|
+
- Turn each hit into a checklist item with its owner and needed permission.
|
|
16
|
+
- Use a supported refactor or link-rewrite tool when one exists; otherwise apply scoped edits.
|
|
17
|
+
- Update indexes and documentation in the same change.
|
|
18
|
+
- When a consumer is outside the granted scope, return the exact required handoff and keep it visible in coverage.
|
|
17
19
|
|
|
18
|
-
##
|
|
20
|
+
## Verify old and new references
|
|
19
21
|
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
Change every item. Use the tool's own governed rename where one exists (a link rewriter, an IDE refactor) over hand edits. If an index or README lists the renamed thing, that index is part of the change, not a follow-up.
|
|
25
|
-
|
|
26
|
-
## 4. Verify: the loud negative (non-negotiable)
|
|
27
|
-
|
|
28
|
-
Re-grep the OLD identifier across every surface. **Expect zero**, except historical records you name explicitly. Then check for dangling references the rename created (unresolved links, 404s, failing imports).
|
|
29
|
-
|
|
30
|
-
Paste the final grep output as proof. If any stray survives, it is not done.
|
|
31
|
-
|
|
32
|
-
## Why the last step is the one that matters
|
|
33
|
-
|
|
34
|
-
Steps 1 to 3 find what you thought of. Step 4 finds what you did not. A rename that "looks complete" and leaves one stray is worse than an unstarted one, because the stray is now believed.
|
|
22
|
+
- Search the old identifier again in every declared scope.
|
|
23
|
+
- Require zero active references, except named compatibility paths or historical records.
|
|
24
|
+
- Verify the new links, imports and consumer behavior with the relevant checker.
|
|
25
|
+
- Return the search command, exit status and any intentional retained matches as evidence.
|
|
@@ -1,165 +1,135 @@
|
|
|
1
|
-
#
|
|
1
|
+
# Lane runner (`aunx cli-run`)
|
|
2
2
|
|
|
3
|
-
`
|
|
3
|
+
When assigning work to a CLI lane, use `aunx cli-run` or the installed `bin/cli-run.mjs`. The runner builds the vendor invocation, reads its terminal event and returns an exit code for the response and any explicit output contract.
|
|
4
4
|
|
|
5
|
-
|
|
5
|
+
Enabled lanes (edit `bin/lanes.json`): {{CLI_RUN_LANES}}.
|
|
6
6
|
|
|
7
|
-
|
|
7
|
+
## Call a lane with a task brief
|
|
8
8
|
|
|
9
9
|
```bash
|
|
10
|
-
|
|
11
|
-
node bin/cli-run.mjs
|
|
10
|
+
aunx cli-run {{EXAMPLE_LANE}} --brief TASK_BRIEF.md --timeout 900
|
|
11
|
+
node bin/cli-run.mjs {{EXAMPLE_LANE}} --brief TASK_BRIEF.md --timeout 900
|
|
12
12
|
```
|
|
13
13
|
|
|
14
|
-
|
|
14
|
+
When `aunx` runs, it uses the packaged runner. Use `--dir <rules-directory>` to select an installed project runner and its lane configuration explicitly. An absent project runner falls back to the package runner.
|
|
15
15
|
|
|
16
|
-
|
|
16
|
+
When requesting a particular route, inspect the lane's current roster and select the model and effort for the job:
|
|
17
17
|
|
|
18
18
|
```bash
|
|
19
|
-
|
|
20
|
-
node bin/cli-run.mjs
|
|
21
|
-
node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high
|
|
22
|
-
node bin/cli-run.mjs codex "<prompt>" --effort auto # medium or high, bounded heuristic
|
|
23
|
-
node bin/cli-run.mjs qwen [--safe-mode] "<prompt>" # qwen-only flag
|
|
19
|
+
aunx cli-run {{EXAMPLE_LANE}} "<prompt>" --model '<model-id>'{{EXAMPLE_EFFORT_FLAGS}}
|
|
20
|
+
node bin/cli-run.mjs {{EXAMPLE_LANE}} "<prompt>" --model '<model-id>'{{EXAMPLE_EFFORT_FLAGS}}
|
|
24
21
|
```
|
|
25
22
|
|
|
26
|
-
|
|
23
|
+
{{EXAMPLE_AUDIT_BLOCK}}
|
|
27
24
|
|
|
28
|
-
|
|
25
|
+
{{QWEN_SAFE_MODE_NOTE}}`--quiet` suppresses human-readable status lines. `--timeout SECS` bounds each call.
|
|
26
|
+
|
|
27
|
+
## Check availability before work depends on it
|
|
29
28
|
|
|
30
29
|
```bash
|
|
31
|
-
|
|
32
|
-
node bin/cli-run.mjs --doctor
|
|
30
|
+
aunx cli-run --doctor
|
|
31
|
+
node bin/cli-run.mjs --doctor
|
|
32
|
+
aunx cli-run --doctor --run
|
|
33
|
+
node bin/cli-run.mjs --doctor --run
|
|
33
34
|
```
|
|
34
35
|
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
A lane that is enabled but not on PATH, or that answers with no deliverable, shows up here before it shows up mid-task.
|
|
38
|
-
|
|
39
|
-
## Why it exists
|
|
40
|
-
|
|
41
|
-
Every agent CLI can exit 0 having produced nothing. The symptom (confident preamble, exit 0, no deliverable) is indistinguishable from a model failure, so it gets blamed on the model. Four wrong diagnoses in one week came from exactly that.
|
|
36
|
+
Use `--doctor` to inspect enabled lanes, binaries and requested defaults. Use `--doctor --run` for a small live prompt on each enabled lane, consuming vendor quota. When no lanes are enabled, the doctor exits 13; choose a supported CLI or use level 1. Verify authentication and loaded project instructions separately from binary presence.
|
|
42
37
|
|
|
43
|
-
|
|
38
|
+
## Verify the result contract
|
|
44
39
|
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
| Lane | Invocation built | Success = |
|
|
48
|
-
|---|---|---|
|
|
49
|
-
| grok | `--output-format json -p` | `stopReason == "end_turn"` and non-empty `text` |
|
|
50
|
-
| codex | `exec --json --color never --skip-git-repo-check -o FILE` | terminal `{"type":"turn.completed"}` and non-empty FILE |
|
|
51
|
-
| agy | `--print-timeout Nm --output-format stream-json -p` | terminal `{"event":"result"}`, `status == "SUCCESS"`, non-empty `response` |
|
|
52
|
-
| hermes | `-z … --usage-file FILE` | its exit code is already honest: 0 response · 1 none · 2 bad args |
|
|
53
|
-
| qwen | `-o json [-m ID] [--safe-mode] -p` | terminal `{"type":"result"}`, `subtype == "success"`, `is_error` false, non-empty `result` not starting with `[API Error:`, and every `stats.models.*.api.totalErrors == 0` |
|
|
40
|
+
When exit 0 returns, the native terminal event indicates completion, the response is non-empty and lane-specific error checks passed. Verify the task's actual acceptance checks separately: a structurally valid refusal can still satisfy that response shape.
|
|
54
41
|
|
|
55
|
-
|
|
42
|
+
When the task requires a file or structured output, add an explicit contract:
|
|
56
43
|
|
|
57
|
-
|
|
44
|
+
```bash
|
|
45
|
+
aunx cli-run {{EXAMPLE_LANE}} "Write out/report.md" --expect-file out/report.md
|
|
46
|
+
node bin/cli-run.mjs {{EXAMPLE_LANE}} "Write out/report.md" --expect-file out/report.md
|
|
47
|
+
aunx cli-run {{EXAMPLE_LANE}} "Return the table as JSON" --expect-json
|
|
48
|
+
node bin/cli-run.mjs {{EXAMPLE_LANE}} "Return the table as JSON" --expect-json
|
|
49
|
+
```
|
|
58
50
|
|
|
59
|
-
|
|
60
|
-
|---|---|---|---|
|
|
61
|
-
| 0 | `ok` | structurally accepted non-empty response, every `--expect-*` contract met | use it; if `refused=N` is above 0, read the problem line |
|
|
62
|
-
| 10 | `empty` | ran and delivered nothing, or a contract was unmet | rerun once, or use another lane |
|
|
63
|
-
| 11 | `no_output` | no output at all | rerun once; check the lane runs on its own |
|
|
64
|
-
| 12 | `timeout` | timed out; the lane and every descendant in its process group were killed | raise `--timeout` or split the brief |
|
|
65
|
-
| 13 | `unavailable` | binary missing, disabled in `lanes.json`, or `lanes.json` malformed | install or enable the lane |
|
|
66
|
-
| 14 | `auth` | the lane's own error says a credential is missing or it is not logged in | set the credential; retrying cannot help |
|
|
67
|
-
| 15 | `quota` | the lane's own error says usage limit, credits or rate limit | switch lanes or wait for the reset |
|
|
68
|
-
| 16 | `rejected` | the upstream rejected the request: an unknown model id, a malformed request | fix the id, flag or request it names |
|
|
69
|
-
| 17 | `refused` | no deliverable, and the lane reports tool calls a hook or deny rule blocked | adjust the rule, or give the lane the tool |
|
|
70
|
-
| 18 | `cut_short` | no trustworthy finish: a missing or non-success terminal event, a lane killed by a signal, output past the 16 MiB buffer, or a nonzero vendor exit nothing above explains (hermes' own exit 2, bad args or an empty response, stays `empty`) | rerun once, then read the lane's stderr |
|
|
71
|
-
| 130 / 143 | `interrupted` | cli-run itself received SIGINT / SIGTERM; the lane's process group was killed first, then the temp dir removed | |
|
|
72
|
-
| 2 | | usage error in cli-run itself | |
|
|
51
|
+
`--expect-file` records existence, size, mtime and content hash before the run. It then requires a non-empty regular file that is new or changed by hash or later mtime. Use a unique per-attempt path when another process might write the same file. `--expect-json` requires parseable JSON. An unmet contract exits 10 with reason `contract_unmet`.
|
|
73
52
|
|
|
74
|
-
|
|
53
|
+
## Read each lane's completion signal
|
|
75
54
|
|
|
76
|
-
|
|
55
|
+
The runner supports these lanes whether or not you selected them.
|
|
77
56
|
|
|
78
|
-
|
|
57
|
+
| Lane | Invocation built | Accepted response |
|
|
58
|
+
|---|---|---|
|
|
59
|
+
| grok | `--output-format json -p` | `stopReason == "end_turn"` and non-empty `text` |
|
|
60
|
+
| codex | `exec --json --color never --skip-git-repo-check -o FILE` | terminal `turn.completed` event and non-empty output file |
|
|
61
|
+
| agy | `--print-timeout Nm --output-format stream-json -p` | terminal `result` event, `status == "SUCCESS"`, non-empty `response` |
|
|
62
|
+
| hermes | `-z` with a usage file | vendor exit 0 and a non-empty response |
|
|
63
|
+
| qwen | `-o json`, optional model and safe-mode flags, `-p` | successful terminal result, no error flag or API-error result, non-empty text and every model's `api.totalErrors == 0` |
|
|
79
64
|
|
|
80
|
-
|
|
81
|
-
cli-run[qwen] exit_nonzero rc=14 class=auth refused=0 0.8s raw=212B route=lane default :: lane exited 1; subtype="error_during_execution": Missing API key ...
|
|
82
|
-
cli-run problem: cli-run[qwen] auth: Missing API key ...
|
|
83
|
-
cli-run fix: set the credential the message above names (its environment variable, or the lane's own login command), then rerun
|
|
84
|
-
```
|
|
65
|
+
When Qwen's error telemetry is absent, the runner refuses the response. These checks distinguish response structure from successful task execution.
|
|
85
66
|
|
|
86
|
-
|
|
67
|
+
## Respond to the exit class
|
|
87
68
|
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
69
|
+
| Code | Class | Action |
|
|
70
|
+
|---|---|---|
|
|
71
|
+
| 0 | `ok` | Verify task acceptance; read the problem line if `refused` is positive |
|
|
72
|
+
| 10 | `empty` | Inspect the missing response or unmet output contract |
|
|
73
|
+
| 11 | `no_output` | Check the lane directly before another attempt |
|
|
74
|
+
| 12 | `timeout` | Diagnose progress, then adjust the bound or split the work |
|
|
75
|
+
| 13 | `unavailable` | Resolve binary presence, lane enablement or malformed configuration |
|
|
76
|
+
| 14 | `auth` | Use the vendor's login or configured credential path |
|
|
77
|
+
| 15 | `quota` | Select another authorized lane or wait for quota renewal |
|
|
78
|
+
| 16 | `rejected` | Correct the model, flag or request named by the vendor |
|
|
79
|
+
| 17 | `refused` | Inspect the denied tool call and route it within authorized permissions |
|
|
80
|
+
| 18 | `cut_short` | Inspect a missing terminal event, signal, output overrun or unexplained nonzero vendor exit |
|
|
81
|
+
| 130 / 143 | `interrupted` | The wrapper received SIGINT / SIGTERM; inspect cleanup before resuming |
|
|
82
|
+
| 2 | usage error | Correct the runner arguments |
|
|
93
83
|
|
|
94
|
-
|
|
84
|
+
When a run fails, read the problem and fix lines and relay the actionable cause. Resolve it within existing authorization; request approval only for a change outside that scope. A nonzero vendor exit remains a failure even if text was produced; `cli_rc` preserves that vendor exit in the local log.
|
|
95
85
|
|
|
96
|
-
|
|
86
|
+
The runner classifies authoritative vendor error fields, with precedence `auth`, `quota`, `rejected`, `refused`, `cut_short`, `empty`. It excludes ordinary model prose from those error signals. A readable result with refused tool calls can remain exit 0; read the reported refusal alongside the result. `refused: null` means the lane supplied no readable refusal signal.
|
|
97
87
|
|
|
98
|
-
|
|
88
|
+
## Record the requested route
|
|
99
89
|
|
|
100
|
-
|
|
101
|
-
node bin/cli-run.mjs codex "<prompt>" --model gpt-6-astra --effort high # this call only
|
|
102
|
-
node bin/cli-run.mjs --doctor # prints what each lane is pinned to
|
|
103
|
-
```
|
|
90
|
+
When a command flag is present, it overrides `bin/lanes.json` defaults. When both are absent, the lane uses its own configuration. Inspect that configuration when the model or effort matters.
|
|
104
91
|
|
|
105
92
|
```json
|
|
106
|
-
{
|
|
107
|
-
"enabled": ["codex", "grok"],
|
|
108
|
-
"defaults": { "codex": { "model": "gpt-6-astra", "effort": "high" } }
|
|
109
|
-
}
|
|
93
|
+
{{EXAMPLE_LANES_JSON}}
|
|
110
94
|
```
|
|
111
95
|
|
|
112
|
-
|
|
96
|
+
Replace the placeholder with a current vendor model ID before using this example. Model and effort values are bounded to the supported safe character set.
|
|
97
|
+
|
|
98
|
+
The runner supports these lanes whether or not you selected them.
|
|
113
99
|
|
|
114
|
-
| Lane | Model |
|
|
100
|
+
| Lane | Model flag | Effort flag |
|
|
115
101
|
|---|---|---|
|
|
116
102
|
| grok | `-m` | `--reasoning-effort` |
|
|
117
103
|
| codex | `-m` | `-c model_reasoning_effort="LEVEL"` |
|
|
118
|
-
| agy | `--model` | `--effort`
|
|
119
|
-
| hermes | `-m` | `--reasoning`
|
|
120
|
-
| qwen | `-m` |
|
|
121
|
-
|
|
122
|
-
Three rules that keep this honest:
|
|
123
|
-
|
|
124
|
-
- **A level `cli-run` does not recognise is not rejected here.** Levels are the vendor's, they change, and guessing the valid set would date this tool. An unknown level is refused by the lane and surfaces as a class (codex reports it as `rejected`, exit 16; a lane with no rejection signal as `cut_short`, exit 18), with the lane's own exit code in the log as `cli_rc` and its stderr on your terminal.
|
|
125
|
-
- **`--effort` on qwen is a usage error, not a silent drop.** A flag that vanishes leaves you believing a route that never ran.
|
|
126
|
-
- **`--effort auto` is bounded.** It uses prompt size, or an audit's changed-file evidence, and resolves only medium or high. An audit is always high. Auto is a heuristic, not a measurement: name `xhigh` explicitly for security-critical or irreversible work.
|
|
127
|
-
- **Values are charset-bounded** (letters, digits, and `. _ : @ / + -`, no leading dash, 64 characters). A model id becomes an argv element and, on codex, part of a TOML value; bounding it is what stops either from being escaped.
|
|
128
|
-
|
|
129
|
-
## Permissions are a separate layer
|
|
130
|
-
|
|
131
|
-
`cli-run` never injects permission flags. Each CLI carries its own config, so every caller gets the same behaviour. Use each vendor's deny-list as the base layer; allow-lists only hold if every binary is enumerable in advance.
|
|
132
|
-
|
|
133
|
-
## Log
|
|
134
|
-
|
|
135
|
-
`~/.ai-orchestrator/cli-run.log.jsonl`, one line per run: lane, verdict, `class` (one of the classes above), rc, the lane's own exit code (`cli_rc`), signal, `refused` (an integer, or `null` when the lane gives no signal), seconds, raw bytes, deliverable bytes, a 12-hex sha256 prefix of the prompt and its length, the route (`model_requested`, `effort_requested`, and `model_source` / `effort_source`, each one of `flag`, `lanes.json` or `lane_default`), auto-sizing evidence (`effort_resolved`, `effort_basis`, `effort_scope`), and `reason`: one of a fixed set of codes (`ok`, `not_json`, `bad_stop_reason`, `empty_text`, `no_terminal_event`, `bad_status`, `api_error_in_result`, `total_errors`, `contract_unmet`, `exit_nonzero`, `timeout`, `killed`, `disabled`, `lanes_json_malformed`, ...). `effort_basis` is exactly one of `explicit`, `prompt_chars`, `audit_floor`, or `none`. Never the prompt text, never a provider-supplied value, never free text: a value the log does not recognise is written as `unknown`. The human-readable detail, which may quote the provider, goes to your terminal only (and nowhere with `--quiet`). "This lane is flaky" becomes a query instead of an argument, and so does "we route audits at high effort".
|
|
136
|
-
|
|
137
|
-
The log records what was **requested**, on every record including a run refused before the lane started. It does not record an actual. Reporting is inconsistent: grok returns a `modelUsage` block naming a model, the other four lanes return nothing of the kind, so an `actual` field would be populated for one lane and empty for four. It would also be a provider-supplied string, and this log holds fixed codes and bounded caller-supplied values only. `model_source: "lane_default"` is the honest way to say this run inherited something invisible from here.
|
|
138
|
-
|
|
139
|
-
## The prompt travels in argv
|
|
140
|
-
|
|
141
|
-
That is each vendor's documented headless shape (`-p`, `exec`). Two consequences: argv is visible to other processes on the machine, so a prompt is never the place for a key; and argv is bounded by the OS (`ARG_MAX`), so a very large brief should be referenced by path inside the prompt rather than pasted whole.
|
|
104
|
+
| agy | `--model` | `--effort` |
|
|
105
|
+
| hermes | `-m` | `--reasoning` |
|
|
106
|
+
| qwen | `-m` | Unsupported; an effort request is a usage error |
|
|
142
107
|
|
|
143
|
-
|
|
108
|
+
When using `--effort auto`, treat its medium/high selection as a bounded heuristic; an audit has a high floor. Use explicit high for builds and xhigh where supported for security-critical or irreversible work. The vendor validates its own effort names and reports unsupported values through the failure class.
|
|
144
109
|
|
|
145
|
-
|
|
110
|
+
## Preserve permissions and secrets
|
|
146
111
|
|
|
147
|
-
|
|
112
|
+
The runner never adds permission flags. Keep each vendor's permissions in its own configuration and route work within the task's granted scope. A denied write becomes a handoff to an authorized writer.
|
|
148
113
|
|
|
149
|
-
|
|
114
|
+
Prompts travel in argv, which other processes may inspect. Never put secrets in a prompt. For large task inputs, give the worker a brief with authorized source paths instead of exceeding the operating system's argument-size limit.
|
|
150
115
|
|
|
151
|
-
|
|
116
|
+
Terminal status and error details are redacted before clipping. Local logs exclude prompt text and provider free text. `--quiet` also suppresses the human-readable details.
|
|
152
117
|
|
|
153
|
-
|
|
118
|
+
## Inspect the local log
|
|
154
119
|
|
|
155
|
-
|
|
120
|
+
Read `~/.ai-orchestrator/cli-run.log.jsonl` for one record per run: lane, verdict, class, return codes, signal, refusal count, timing, byte counts, prompt hash and length, requested model and effort, source of those requests, auto-effort evidence and a fixed reason code.
|
|
156
121
|
|
|
157
|
-
|
|
122
|
+
Use `model_source: "lane_default"` to identify an inherited route. These fields record requested settings; verifying the vendor's actual model requires vendor evidence. Keep task success grounded in the acceptance checks.
|
|
158
123
|
|
|
159
|
-
##
|
|
124
|
+
## Configuration and process boundaries
|
|
160
125
|
|
|
161
|
-
|
|
126
|
+
- When `lanes.json` is absent, all supported lanes are enabled with inherited defaults. When it is malformed or unreadable, the runner refuses every lane with exit 13 until corrected.
|
|
127
|
+
- When a lane dies by signal, discard its partial response and handle `cut_short` exit 18.
|
|
128
|
+
- On POSIX, timeout, output overrun and catchable interruption kill the lane's process group. A child that creates its own session can escape that boundary; use a service-level process boundary where needed.
|
|
129
|
+
- On Windows, lanes use a direct executable or a resolved Node shim, never `cmd.exe`. Windows termination behavior differs from POSIX signal cleanup.
|
|
130
|
+
- When the wrapper receives uncatchable SIGKILL, use a supervisor such as systemd with `KillMode=control-group` to clean up its process tree.
|
|
131
|
+
- Vendor output uses streaming UTF-8 decoding and a 16 MiB cap counted in bytes.
|
|
162
132
|
|
|
163
|
-
##
|
|
133
|
+
## Choose the lane
|
|
164
134
|
|
|
165
|
-
|
|
135
|
+
When selecting a lane, apply `ROUTING.md` and `DELEGATION_MATRIX.md`, optionally starting with `aunx route "<task>"`. The runner executes the lane named by the caller. Its response checks work independently of optional companion software.
|
|
@@ -1,7 +1,13 @@
|
|
|
1
|
-
# DELEGATION_MATRIX.md:
|
|
1
|
+
# DELEGATION_MATRIX.md: choose a lane for the task
|
|
2
2
|
|
|
3
3
|
Generated {{DATE}} from the AIs you said you have: `{{AI_IDS}}`.
|
|
4
4
|
|
|
5
|
+
{{STACK_TABLE}}
|
|
6
|
+
|
|
7
|
+
{{STACK_FALLBACK_NOTE}}
|
|
8
|
+
|
|
9
|
+
{{STACK_GAPS}}
|
|
10
|
+
|
|
5
11
|
## Your lanes
|
|
6
12
|
|
|
7
13
|
{{LANES_TABLE}}
|
|
@@ -10,7 +16,7 @@ Generated {{DATE}} from the AIs you said you have: `{{AI_IDS}}`.
|
|
|
10
16
|
|
|
11
17
|
{{PLAN_GUIDANCE}}
|
|
12
18
|
|
|
13
|
-
##
|
|
19
|
+
## Match the task to an available lane
|
|
14
20
|
|
|
15
21
|
{{TASK_LANES_TABLE}}
|
|
16
22
|
|
|
@@ -24,4 +30,4 @@ Generated {{DATE}} from the AIs you said you have: `{{AI_IDS}}`.
|
|
|
24
30
|
|
|
25
31
|
## Privacy gate
|
|
26
32
|
|
|
27
|
-
|
|
33
|
+
Before dispatching protected data, check the project's named allowed and barred lanes. Never send private notes, client data or personal records to an unapproved third-party bulk or fan-out lane. When the required permission or tool is absent, keep the data local and return the needed decision.
|
|
@@ -5,9 +5,9 @@ Written at level 2 and above, on top of `common/` and `beginner/`.
|
|
|
5
5
|
| File | What it adds |
|
|
6
6
|
|---|---|
|
|
7
7
|
| `ROUTING.md` | the multi-lane decision tree; supersedes `ORCHESTRATOR.md` when present |
|
|
8
|
-
| `TIERS.md` | capability tiers, the three cost levers,
|
|
8
|
+
| `TIERS.md` | capability tiers, the three cost levers, job fit, model availability and effort selection |
|
|
9
9
|
| `DELEGATION_MATRIX.md` | task → lane → pick, generated from the user's selection |
|
|
10
|
-
| `RESEARCH_TRIAGE.md` |
|
|
11
|
-
| `CLI-RUN.md` | how `bin/cli-run.mjs`
|
|
10
|
+
| `RESEARCH_TRIAGE.md` | selected engines in parallel, one source-checking writer |
|
|
11
|
+
| `CLI-RUN.md` | how `aunx cli-run` and `bin/cli-run.mjs` verify each response |
|
|
12
12
|
|
|
13
13
|
`bin/cli-run.mjs` and `bin/lanes.json` are written by the installer from `bin/cli-run.mjs` in this repo and the user's selection; they are not templates.
|
|
@@ -1,30 +1,38 @@
|
|
|
1
|
-
# RESEARCH_TRIAGE.md:
|
|
1
|
+
# RESEARCH_TRIAGE.md: parallel research with source verification
|
|
2
2
|
|
|
3
|
-
{{RESEARCH_SELECTION_ADVICE}} Then
|
|
3
|
+
When a question needs outside evidence, {{RESEARCH_SELECTION_ADVICE}} Then open the primary sources that support decisions.
|
|
4
4
|
|
|
5
|
-
|
|
5
|
+
Use the selected `cli-run` lanes: {{CLI_RUN_LANES}} ({{RESEARCH_ENGINES}} research engine(s) below). When a needed lane is absent, use an available authorized research tool and name the resulting coverage limit.
|
|
6
6
|
|
|
7
|
-
##
|
|
7
|
+
## Assign the research roles
|
|
8
8
|
|
|
9
9
|
| Role | Typical lane | Job |
|
|
10
10
|
|---|---|---|
|
|
11
|
-
{{RESEARCH_ROLES}}
|
|
11
|
+
{{RESEARCH_ROLES}}
|
|
12
12
|
|
|
13
|
-
|
|
13
|
+
Give each engine a task brief through `--brief`, with the same context file, bounded questions, source standard and stopping condition. When a run returns no usable result, record that engine as unavailable and continue independent source checks.
|
|
14
14
|
|
|
15
|
-
##
|
|
15
|
+
## Run the selected engines
|
|
16
16
|
|
|
17
17
|
```bash
|
|
18
|
-
BRIEF=research/brief.md
|
|
18
|
+
BRIEF=research/brief.md
|
|
19
19
|
{{RESEARCH_RUN}}
|
|
20
20
|
```
|
|
21
21
|
|
|
22
|
-
|
|
22
|
+
When using the installed script directly, `aunx cli-run` accepts the same runner arguments as `node bin/cli-run.mjs`.
|
|
23
23
|
|
|
24
|
-
|
|
24
|
+
When outputs arrive, have one writer inspect the primary sources carrying each decision and produce a dated synthesis:
|
|
25
25
|
|
|
26
|
-
-
|
|
27
|
-
-
|
|
28
|
-
-
|
|
29
|
-
-
|
|
30
|
-
|
|
26
|
+
- **CONFIRMED:** verified against the primary source, with independent corroboration where the question needs it.
|
|
27
|
+
- **DISAGREEMENT:** preserve conflicting readings and identify the evidence that would resolve them.
|
|
28
|
+
- **REPORTED:** attribute a source's statement to that source.
|
|
29
|
+
- **UNVERIFIED:** name the missing evidence or access.
|
|
30
|
+
|
|
31
|
+
## Verify claims and coverage
|
|
32
|
+
|
|
33
|
+
- When engines agree, verify the shared premise against source.
|
|
34
|
+
- When engines disagree, preserve both claims until evidence resolves them.
|
|
35
|
+
- When a report gives a number, open its source or recompute it with an available tool.
|
|
36
|
+
- When testing a research method, use a labelled false-premise fixture and confirm the method rejects it before relying on it.
|
|
37
|
+
- When recording results, use one writer and return claim dispositions with citations.
|
|
38
|
+
- When Context7 or another companion is absent, use official documentation, source and the project's own runtime checks.
|
|
@@ -1,8 +1,16 @@
|
|
|
1
|
-
# ROUTING.md:
|
|
1
|
+
# ROUTING.md: model router for your available tools
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
Main agent: **{{PRIMARY_NAME}}**. A **lane** is an AI tool or model the main agent can hand work to. A **tier** describes a model's capability and cost: planning model, working model or cheap model.
|
|
4
4
|
|
|
5
|
-
|
|
5
|
+
When a task arrives, choose its route from the live tools and these rules. `aunx route "<task>"` prints a deterministic keyword suggestion; verify that suggestion against the task's scope, required tools and stakes before dispatching.
|
|
6
|
+
|
|
7
|
+
{{STACK_TABLE}}
|
|
8
|
+
|
|
9
|
+
{{STACK_FALLBACK_NOTE}}
|
|
10
|
+
|
|
11
|
+
{{STACK_GAPS}}
|
|
12
|
+
|
|
13
|
+
## Your lanes
|
|
6
14
|
|
|
7
15
|
{{LANES_TABLE}}
|
|
8
16
|
|
|
@@ -10,76 +18,71 @@ Your lanes:
|
|
|
10
18
|
|
|
11
19
|
{{PLAN_GUIDANCE}}
|
|
12
20
|
|
|
13
|
-
|
|
21
|
+
When choosing a billing route, distinguish **subscription lanes** (usage covered by a subscription, subject to its limits) from **pay-per-token lanes** (metered API usage). When data must stay on the machine, choose a verified **local lane** and check its network behavior.
|
|
14
22
|
|
|
15
|
-
|
|
23
|
+
When a cheaper eligible route can complete and verify the task, select it. When verification fails, complexity emerges or tool reach is insufficient, escalate with a named reason.
|
|
16
24
|
|
|
17
25
|
## Decision tree (first match wins)
|
|
18
26
|
|
|
19
|
-
0. **
|
|
27
|
+
0. **An external lane offers useful reach or capacity:** check `DELEGATION_MATRIX.md` and call the selected CLI through `aunx cli-run` or the installed `bin/cli-run.mjs`.
|
|
20
28
|
{{LANE_STEP0}}
|
|
21
|
-
1. **Bulk
|
|
22
|
-
1a. **
|
|
23
|
-
2. **
|
|
24
|
-
3. **
|
|
25
|
-
3a. **
|
|
26
|
-
3b. **
|
|
27
|
-
4. **
|
|
29
|
+
1. **Bulk or mechanical work:** classify, tag, extract, rename or reformat -> cheap model tier / {{BULK_ROLE}}{{BULK_LANE}}.{{FAN_OUT_ADVICE}}
|
|
30
|
+
1a. **Read or digest many files:** -> {{READER_ROLE}}; return facts, quotes or an index within the brief's scope.
|
|
31
|
+
2. **Current data is required:** -> {{LIVE_LANE}} working model tier with live tools.
|
|
32
|
+
3. **Review code without changing it:** -> {{REVIEW_ROLE}}, working model tier. For security-critical scope -> {{ATTACK_LANE}} with appropriate effort.
|
|
33
|
+
3a. **A reviewer or scanner has returned findings:** -> {{FINDING_ROLE}}; reproduce each claim before repair.
|
|
34
|
+
3b. **Check a task's definition of done:** -> {{DONE_ROLE}}; probe the named artifact and return MET, NOT_MET or UNVERIFIABLE.
|
|
35
|
+
4. **Ambiguity, architecture or an unknown cause:** -> planning model tier / {{PLANNER_ROLE}}. Return a concrete plan for execution.
|
|
28
36
|
{{DECISION_RULE5}}
|
|
29
37
|
|
|
30
38
|
{{WHO_BUILDS}}
|
|
31
39
|
|
|
32
|
-
##
|
|
40
|
+
## Build process and ownership
|
|
33
41
|
|
|
34
|
-
|
|
35
|
-
|---|---|
|
|
36
|
-
| 0 Route | live probe for access; `cli-run` lanes are $0 and uncapped |
|
|
37
|
-
| 1 Map | the orchestrator sweeps{{STAGE1_LANES}} |
|
|
38
|
-
| 2 Judge | deep tier, on the finished map: one named weak spot and one gap in the request |
|
|
39
|
-
| 3 Build | the orchestrator, against the installed dependency's source |
|
|
40
|
-
| 4 Scan | secret + static + dependency scanners, diff-scoped, refuses by default |
|
|
41
|
-
| 5 Challenge | security-shaped diff → {{ATTACK_LANE}}. Architecture-shaped → deep tier, build against plan. Never both |
|
|
42
|
-
| 5a Verify findings | finding-verifier, a different model family where you have one: CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Only CONFIRMED earns a repair |
|
|
43
|
-
| 5b Ship | rollback id recorded, explicit human yes |
|
|
44
|
-
| 6 Verify | real test, negative test seen red, old identifier re-grepped to zero |
|
|
45
|
-
| 7 Record | one end-to-end doc, tracker Done with evidence, plan doc deleted |
|
|
42
|
+
When a task builds or changes a system, run `protocols/build-protocol.md`.
|
|
46
43
|
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
44
|
+
| Step | Action |
|
|
45
|
+
|---|---|
|
|
46
|
+
| Frame and probe | Quote the ask, freeze acceptance checks and verify required availability |
|
|
47
|
+
| Ordering | Dependencies first, invalidators early, deterministic checks before judgment, irreversible actions last |
|
|
48
|
+
| Research | Vet sources, inspect current interfaces and spike risky assumptions |
|
|
49
|
+
| Context file | Map affected surfaces once; every brief reads the same context file |
|
|
50
|
+
| Assign | Choose each section's lane, model and effort by live capability and job fit |
|
|
51
|
+
| Build | Execute the approved scope; hand refused writes to an authorized writer |
|
|
52
|
+
| Split and merge | Keep the whole scope in each brief; merge sections and name conflicts |
|
|
53
|
+
| Audit | One pass on the merged artifact plus a companion consult asking scope versus ask |
|
|
54
|
+
| Ship | Record rollback, confirm authorization, replay checks and verify the change in use |
|
|
55
|
+
| Record | Return coverage and evidence, update documentation and name what watches it |
|
|
54
56
|
|
|
55
|
-
|
|
57
|
+
When work runs in the background, check liveness and output growth every five minutes. Two checks without growth call for diagnosis and a report. When an audit finding is confirmed, assign its fix to a non-author and show the regression failing before the fix. Verify the fix; no second audit pass.
|
|
56
58
|
|
|
57
|
-
##
|
|
59
|
+
## Tools and fallbacks
|
|
58
60
|
|
|
59
|
-
|
|
61
|
+
- When reporting consequential arithmetic or code equivalence, use a computing tool (`protocols/numbers-and-logic.md`). codecalc: {{CODECALC_STATUS}}. When absent, use the local runtime, test suite or spreadsheet.{{METERED_CITATION_NOTE}}
|
|
62
|
+
- When writing durable records, search first, update the index and keep one writer (`protocols/memory-and-record.md`). obsidian-tc: {{OBSIDIAN_TC_STATUS}}. When absent, use project files, search and version control.
|
|
63
|
+
- When using a changing library or API, read current documentation and verify behavior (`protocols/docs-then-prove.md`). Context7: {{CONTEXT7_STATUS}}. When absent, read official docs or installed source; use the local runtime when codecalc is absent.
|
|
60
64
|
|
|
61
|
-
##
|
|
65
|
+
## Choose effort and verify the route
|
|
62
66
|
|
|
63
67
|
{{PLAN_BIG_LINE}}{{INLINE_THRESHOLD_NOTE}}
|
|
64
|
-
-
|
|
65
|
-
-
|
|
66
|
-
-
|
|
67
|
-
-
|
|
68
|
-
-
|
|
69
|
-
-
|
|
70
|
-
- **Pin the route when it matters:** a lane with no `--model`/`--effort` and no `defaults` entry in `bin/lanes.json` runs on its own config, which may be nothing like what this file describes. `cli-run --doctor` prints what each lane is pinned to, and every run logs the value requested and where it came from.
|
|
68
|
+
- When a lookup is sufficient, use the cheap model tier; when the task needs judgment, choose the working or planning model tier by evidence.
|
|
69
|
+
- When an attempt fails, identify the failure class before retrying; change the route or resolve the cause explicitly.
|
|
70
|
+
- When delegating, pass the context file and task brief with scoped source references and acceptance checks.
|
|
71
|
+
- When role, complexity or stakes change, reassess model, effort and reviewer together. See `TIERS.md`.
|
|
72
|
+
- Before each build, inspect the lane's live roster and any configured pin. Flag stale pins and record the model and effort selected for this task.
|
|
73
|
+
- When the route matters, set `--model` and `--effort` or verify `bin/lanes.json` defaults. Run `aunx cli-run --doctor` to inspect requested defaults; an unpinned lane uses its own configuration.
|
|
71
74
|
|
|
72
75
|
## Example routings
|
|
73
76
|
|
|
74
77
|
| Task | Route |
|
|
75
78
|
|---|---|
|
|
76
|
-
|
|
|
77
|
-
|
|
|
79
|
+
| Design the architecture for a service | {{PLANNER_ROLE}}, planning model tier |
|
|
80
|
+
| Review this service for bugs | {{REVIEW_ROLE}}, working model tier |
|
|
78
81
|
{{ADD_ENDPOINT_ROW}}
|
|
79
|
-
|
|
|
80
|
-
|
|
|
81
|
-
|
|
|
82
|
-
|
|
|
83
|
-
|
|
|
82
|
+
| Find why this silently drops rows | {{PLANNER_ROLE}}, then a scoped build |
|
|
83
|
+
| Summarize similar notes into one index | {{BULK_ROLE}}, cheap model tier |
|
|
84
|
+
| Read every file and extract mentions of a topic | {{READER_ROLE}}, cheap model tier |
|
|
85
|
+
| Verify the audit's findings | {{FINDING_ROLE}} before repairs |
|
|
86
|
+
| Check whether the stated definition of done holds | {{DONE_ROLE}} |
|
|
84
87
|
{{LANE_EXAMPLES}}
|
|
85
88
|
{{ROUTE_GATE_SECTION}}
|