model-orchestrator 0.1.15 → 0.1.17
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +27 -1
- package/README.md +29 -14
- package/bin/cli-run.mjs +22 -7
- package/bin/cli.js +3 -3
- package/docs/README.md +1 -1
- package/docs/audit-brief.md +11 -0
- package/docs/catalog.md +1 -1
- package/docs/part-1-beginner.md +4 -4
- package/docs/part-2-intermediate.md +5 -5
- package/llms.txt +1 -1
- package/package.json +1 -1
- package/src/catalog.js +1 -1
- package/src/detect.js +28 -9
- package/src/install.js +31 -10
- package/templates/agents/agy/README.md +1 -1
- package/templates/agents/agy/builder.md +1 -1
- package/templates/agents/agy/finding-verifier.md +7 -1
- package/templates/agents/claude-code/README.md +2 -2
- package/templates/agents/claude-code/code-reviewer.md +9 -2
- package/templates/agents/claude-code/finding-verifier.md +9 -2
- package/templates/agents/snippets/chat.md +2 -2
- package/templates/agents/snippets/claude-code.md +2 -2
- package/templates/agents/snippets/generic.md +1 -1
- package/templates/agents/snippets/route-gate.mjs +1 -1
- package/templates/agents/snippets/route-metrics.mjs +356 -0
- package/templates/agents/snippets/settings.hooks.snippet.json +44 -0
- package/templates/beginner/ORCHESTRATOR.md +2 -2
- package/templates/common/README.md +1 -1
- package/templates/common/protocols/build-protocol.md +10 -10
- package/templates/common/protocols/deep-research.md +2 -2
- package/templates/common/protocols/gap-analysis.md +1 -1
- package/templates/common/protocols/propagate.md +2 -2
- package/templates/intermediate/CLI-RUN.md +3 -3
- package/templates/intermediate/DELEGATION_MATRIX.md +1 -1
- package/templates/intermediate/ROUTING.md +4 -4
- package/templates/intermediate/TIERS.md +12 -8
- package/templates/tools/codecalc/CODECALC.md +1 -1
- package/templates/tools/obsidian-tc/OBSIDIAN-TC.md +3 -3
|
@@ -31,11 +31,11 @@ Agreement is weak evidence. Disagreement is the signal.
|
|
|
31
31
|
|
|
32
32
|
## Level 1: one agent
|
|
33
33
|
|
|
34
|
-
You still get the shape. Run PLAN as its own turn and inspect it before spending anything. Run the sweep. Then run a **fresh-context
|
|
34
|
+
You still get the shape. Run PLAN as its own turn and inspect it before spending anything. Run the sweep. Then run a **fresh-context second-opinion turn** with a brief that says "question the premise; list what this report would get wrong if its sources were stale". Plant one deliberately wrong figure in the brief and see whether it corrects it: if it does not, its confirmations are worth less than they look. Mark every claim.
|
|
35
35
|
|
|
36
36
|
## Level 2 and up: three engines, one triager
|
|
37
37
|
|
|
38
|
-
Fan out the same PLAN to three different model families through their CLIs (a web-sweep lane,
|
|
38
|
+
Fan out the same PLAN to three different model families through their CLIs (a web-sweep lane, a second-opinion-read lane, a live-data lane). Run them through `cli-run` so a run that produced nothing is caught as `rc=10` rather than read as an empty finding. The orchestrator triages: it opens the primary sources itself, marks each claim, and writes the brief. Only the orchestrator writes the durable record; every other engine proposes.
|
|
39
39
|
|
|
40
40
|
Known failure shape: one engine will return confident unsourced numerics and claim full coverage. Downgrade those to hypothesis. The engines that report their own gaps honestly are the ones to weight.
|
|
41
41
|
|
|
@@ -14,7 +14,7 @@ Verification asks "is what I did correct?". Gap analysis asks "what did I not do
|
|
|
14
14
|
## Who runs it
|
|
15
15
|
|
|
16
16
|
- **Level 1 (one agent):** the same agent, in a fresh turn, with a brief that says "you are looking for what is missing; do not re-verify what is present". Fresh context matters more than a different model.
|
|
17
|
-
- **Level 2 and up:** a **different model family** reading the same artifact. Disagreement between two families is the cheapest available signal that something is soft. The
|
|
17
|
+
- **Level 2 and up:** a **different model family** reading the same artifact. Disagreement between two families is the cheapest available signal that something is soft. The second-opinion coder lane (read-only mode) is the natural fit.
|
|
18
18
|
- **Level 3:** make it recurring. A weekly audit job enumerates live state (lanes, jobs, services, model lists), diffs it against the plan, and files a report. It catches the dead lane and the silently renamed model nobody noticed.
|
|
19
19
|
|
|
20
20
|
## The second half: analyze, compare, suggest
|
|
@@ -1,10 +1,10 @@
|
|
|
1
1
|
# Propagate: change completeness
|
|
2
2
|
|
|
3
|
-
**A rename is a refactor, not a single-file edit.** Any change to a name, term, path, slug, schema field, routing rule or shared convention
|
|
3
|
+
**A rename is a refactor, not a single-file edit.** Any change to a name, term, path, slug, schema field, routing rule or shared convention reaches everything that uses it, and the goal is zero silent strays.
|
|
4
4
|
|
|
5
5
|
This is retrieval work. It stays with the orchestrator (or a cheap worker for the grep sweep). It never goes to the deep tier: a judgment model re-deriving a file list is the most expensive routing mistake there is.
|
|
6
6
|
|
|
7
|
-
## 1. Map
|
|
7
|
+
## 1. Map everything it touches (before editing anything)
|
|
8
8
|
|
|
9
9
|
- **Docs and notes:** backlinks to the thing being renamed; literal search for the old term and its link forms. With obsidian-tc: `get_backlinks`, `search_text`, then `find_unresolved_links` after the change (`protocols/memory-and-record.md`).
|
|
10
10
|
- **Memory / instructions:** grep every instructions file your agents read (`CLAUDE.md`, `AGENTS.md`, `GEMINI.md`, `QWEN.md`, custom instructions) and any memory store.
|
|
@@ -68,7 +68,7 @@ qwen is the lane whose own success flags lie: an upstream 400 comes back as exit
|
|
|
68
68
|
|
|
69
69
|
## The route: which model, and how hard it thinks
|
|
70
70
|
|
|
71
|
-
A lane you do not pin runs on **its own config file**, which this tool cannot see. That is the quiet failure this section exists for: a CLI configured months ago at `reasoning_effort = "low"` keeps auditing at low effort while your routing docs describe
|
|
71
|
+
A lane you do not pin runs on **its own config file**, which this tool cannot see. That is the quiet failure this section exists for: a CLI configured months ago at `reasoning_effort = "low"` keeps auditing at low effort while your routing docs describe a second-opinion pass, and nothing anywhere says so.
|
|
72
72
|
|
|
73
73
|
Pin it per call, or per lane:
|
|
74
74
|
|
|
@@ -114,9 +114,9 @@ The log records what was **requested**, on every record including a run refused
|
|
|
114
114
|
|
|
115
115
|
That is each vendor's documented headless shape (`-p`, `exec`). Two consequences: argv is visible to other processes on the machine, so a prompt is never the place for a key; and argv is bounded by the OS (`ARG_MAX`), so a very large brief should be referenced by path inside the prompt rather than pasted whole.
|
|
116
116
|
|
|
117
|
-
## lanes.json
|
|
117
|
+
## lanes.json refuses by default
|
|
118
118
|
|
|
119
|
-
Absent: every lane enabled, nothing pinned. Present but malformed or unreadable: every lane refused (exit 13) until it is fixed. A half-written config never re-enables a lane the installer disabled. `defaults` is optional and held to the same standard: a malformed entry, an unknown lane, an unknown key, a value outside the charset, or an effort pinned on a lane that has no reasoning flag all
|
|
119
|
+
Absent: every lane enabled, nothing pinned. Present but malformed or unreadable: every lane refused (exit 13) until it is fixed. A half-written config never re-enables a lane the installer disabled. `defaults` is optional and held to the same standard: a malformed entry, an unknown lane, an unknown key, a value outside the charset, or an effort pinned on a lane that has no reasoning flag all make the whole file refuse by default rather than being skipped quietly.
|
|
120
120
|
|
|
121
121
|
## A killed lane is not a deliverable
|
|
122
122
|
|
|
@@ -15,7 +15,7 @@ Generated {{DATE}} from the AIs you said you have: `{{AI_IDS}}`.
|
|
|
15
15
|
| Many independent items each needing its own agent turn | a concurrent fan-out lane | one call, N children, on a subscription |
|
|
16
16
|
| Live web or social reads | the live-data CLI | subscription-covered; the same search on the API bills per call |
|
|
17
17
|
| Code review, no changes | standard tier, or the second-coder CLI | a different model family catches what one misses |
|
|
18
|
-
|
|
|
18
|
+
| Second-opinion audit of a security-shaped diff | the second-coder CLI in read-only audit mode | Claude writes, a second family challenges, the orchestrator reproduces |
|
|
19
19
|
| Deep architecture / planning | deep tier | expensive to get wrong |
|
|
20
20
|
| Well-specified execution | the orchestrator | execution does not need the top tier |
|
|
21
21
|
| Long-document analysis | the largest-context lane, or caching on the primary | window size vs re-query cost |
|
|
@@ -31,10 +31,10 @@ Rule of thumb: never spend a frontier token on a task a cheap tier finishes corr
|
|
|
31
31
|
|---|---|
|
|
32
32
|
| 0 Route | live probe for access; `cli-run` lanes are $0 and uncapped |
|
|
33
33
|
| 1 Map | the orchestrator sweeps{{STAGE1_LANES}} |
|
|
34
|
-
| 2 Judge | deep tier, on the finished map:
|
|
34
|
+
| 2 Judge | deep tier, on the finished map: one named weak spot and one gap in the request |
|
|
35
35
|
| 3 Build | the orchestrator, against the installed dependency's source |
|
|
36
|
-
| 4 Scan | secret + static + dependency scanners, diff-scoped,
|
|
37
|
-
| 5
|
|
36
|
+
| 4 Scan | secret + static + dependency scanners, diff-scoped, refuses by default |
|
|
37
|
+
| 5 Challenge | security-shaped diff → {{ATTACK_LANE}}. Architecture-shaped → deep tier, build against plan. Never both |
|
|
38
38
|
| 5a Verify findings | finding-verifier, a different model family where you have one: CONFIRMED, NOT_REPRODUCED or INCONCLUSIVE per finding. Only CONFIRMED earns a repair |
|
|
39
39
|
| 5b Ship | rollback id recorded, explicit human yes |
|
|
40
40
|
| 6 Verify | real test, negative test seen red, old identifier re-grepped to zero |
|
|
@@ -58,7 +58,7 @@ One writer per run; every other lane proposes. Search before writing, index in t
|
|
|
58
58
|
- **Long context:** mechanical digestion → fast tier in chunks; judgment over a long input → standard tier.
|
|
59
59
|
- **Token discipline on every delegation:** pass only the context the delegate needs, never the conversation.
|
|
60
60
|
- **Effort per agent:** deep xhigh, review, verification and build high, live research medium, bulk low.
|
|
61
|
-
- **Three inputs, not one:** role picks the agent, complexity moves the effort,
|
|
61
|
+
- **Three inputs, not one:** role picks the agent, complexity moves the effort, stakes move the tier and who reads it. A one-line auth change is simple and high-stakes at once, and the stakes decide. See `TIERS.md`.
|
|
62
62
|
- **Pin the route when it matters:** a lane with no `--model`/`--effort` and no `defaults` entry in `bin/lanes.json` runs on its own config, which may be nothing like what this file describes. `cli-run --doctor` prints what each lane is pinned to, and every run logs the value requested and where it came from.
|
|
63
63
|
|
|
64
64
|
## Example routings
|
|
@@ -47,21 +47,25 @@ reasoning than the reviewer judging its output.** When the plan is airtight the
|
|
|
47
47
|
spec is carrying the thinking, so builder drops to medium. When the plan is
|
|
48
48
|
vague, fix the plan; do not buy reasoning to paper over it.
|
|
49
49
|
|
|
50
|
-
**
|
|
50
|
+
**Stakes move the tier and the reader, never just the effort.** These four are
|
|
51
51
|
the ones worth naming, because their failures are not recoverable by editing the
|
|
52
52
|
code afterwards.
|
|
53
53
|
|
|
54
|
-
|
|
54
|
+
Stakes means what a mistake would cost: a security hole, leaked personal data,
|
|
55
|
+
lost data, or something you can't undo. Most tasks are low-stakes and route
|
|
56
|
+
normally.
|
|
57
|
+
|
|
58
|
+
| Stakes | Present when the change touches | What it buys |
|
|
55
59
|
|---|---|---|
|
|
56
|
-
| security | auth, tokens, sessions, routes, untrusted input | the
|
|
60
|
+
| security | auth, tokens, sessions, routes, untrusted input | the challenge pass, ideally a different model family |
|
|
57
61
|
| privacy | personal data, anything leaving the machine | the local lane, and a named check on what is sent |
|
|
58
62
|
| data loss | deletion, bulk mutation, migrations, overwrites | a reviewed rollback path before the change is written |
|
|
59
63
|
| irreversible | publishing, sending, rotating, anything with an audience | a human yes at Stage 5b, never an agent's |
|
|
60
64
|
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
synonym for difficulty: a one-line change to an auth check is simple and
|
|
64
|
-
high-
|
|
65
|
+
High stakes raise code-reviewer to xhigh, and a security-shaped diff goes to
|
|
66
|
+
the challenge lane rather than to a second read by the same family. Stakes are
|
|
67
|
+
not a synonym for difficulty: a one-line change to an auth check is simple and
|
|
68
|
+
high-stakes at the same time, and it is the stakes that decide the route.
|
|
65
69
|
|
|
66
70
|
**Reserve the top of the ladder for evidence.** xhigh and the escalation tier are
|
|
67
71
|
bought with a named reason: a reproduced failure, a checkpoint that came back
|
|
@@ -70,7 +74,7 @@ task, not an escalation.
|
|
|
70
74
|
|
|
71
75
|
## Why split tiers: robustness first, cost second
|
|
72
76
|
|
|
73
|
-
The split produces better work. The deep tier steers every build twice, and what it steers is **judgment, never retrieval**: the orchestrator sweeps
|
|
77
|
+
The split produces better work. The deep tier steers every build twice, and what it steers is **judgment, never retrieval**: the orchestrator sweeps everything it touches itself and hands the deep tier a finished map. Paying deep-tier rates for a file list is the most expensive routing mistake available.
|
|
74
78
|
|
|
75
79
|
Against a baseline of "standard tier with no consults", default checkpoints are a spend increase. That is the accepted trade, not a saving to claim.
|
|
76
80
|
|
|
@@ -36,7 +36,7 @@ The tools cannot help a model that never reaches for them. codecalc ships `SKILL
|
|
|
36
36
|
|
|
37
37
|
## What it is not
|
|
38
38
|
|
|
39
|
-
Not a cloud sandbox for multi-tenant loads, not a replacement for a vendor's built-in interpreter when zero setup matters more than measurement.
|
|
39
|
+
Not a cloud sandbox for multi-tenant loads, not a replacement for a vendor's built-in interpreter when zero setup matters more than measurement. It assumes a single-operator, local, stdio setup. It earns its keep when the correctness of a claim, not "it ran", is the point.
|
|
40
40
|
|
|
41
41
|
## On a box (level 3)
|
|
42
42
|
|
|
@@ -11,7 +11,7 @@ A durable, searchable, governed store that the protocols can call by name:
|
|
|
11
11
|
| Need in the protocols | obsidian-tc tool |
|
|
12
12
|
|---|---|
|
|
13
13
|
| find what exists before writing (deep research dedupe, gap analysis) | `semantic_search`, `search_text`, `search_regex` |
|
|
14
|
-
| map a rename
|
|
14
|
+
| map everything a rename touches (propagate) | `get_backlinks`, `find_unresolved_links`, `rewrite_link` |
|
|
15
15
|
| record the end-to-end doc (build Stage 7) | `write_note` (compare-and-swap, confirmation on overwrite), `patch_note`, `append_note` |
|
|
16
16
|
| keep inferred content honest | `write_note` with `provenance: "agent_synthesis"` runs a poison scan before the write lands |
|
|
17
17
|
| keep a shared vault safe for several agents | JWT scopes, per-vault folder ACLs, a read-only kill switch, human-in-the-loop tokens |
|
|
@@ -56,9 +56,9 @@ Merge the block; do not replace the file.
|
|
|
56
56
|
|
|
57
57
|
## Security posture, read before a second agent touches it
|
|
58
58
|
|
|
59
|
-
Zero-config mode boots with **auth off and no folder ACL**: anything that can reach the server has the same authority as raw filesystem access to the vault. That is acceptable only because the surface is local-only (the config
|
|
59
|
+
Zero-config mode boots with **auth off and no folder ACL**: anything that can reach the server has the same authority as raw filesystem access to the vault. That is acceptable only because the surface is local-only (the config refuses by default if you enable HTTP on a non-loopback host with auth off, and a DNS-rebinding guard protects loopback). Before exposing it to partially-trusted, remote or multi-agent callers, turn on `auth.mode: "jwt"` and set `acl.readPaths` / `writePaths` / `deletePaths` in the config file. Upstream `SECURITY.md` has the security notes and a private disclosure path.
|
|
60
60
|
|
|
61
|
-
Track record worth knowing: an independent code audit of v1.8.1 (July 2026) found three security-relevant gaps (an ACL
|
|
61
|
+
Track record worth knowing: an independent code audit of v1.8.1 (July 2026) found three security-relevant gaps (an ACL bypass that let enumeration tools skip its refuse-by-default rule, a compare-and-swap bypass through `upsert`, a poison-eligibility gap in preference extraction). All three were fixed upstream before they were filed; verified against the v1.25.0 source on 2026-09-03.
|
|
62
62
|
|
|
63
63
|
## Level 3
|
|
64
64
|
|