model-orchestrator 0.1.14 → 0.1.16

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (36) hide show
  1. package/AGENTS.md +26 -0
  2. package/CHANGELOG.md +43 -2
  3. package/README.md +52 -8
  4. package/bin/cli-run.mjs +22 -7
  5. package/bin/cli.js +4 -4
  6. package/docs/audit-brief.md +32 -0
  7. package/docs/part-1-beginner.md +1 -1
  8. package/docs/part-2-intermediate.md +1 -1
  9. package/llms.txt +27 -0
  10. package/package.json +28 -4
  11. package/src/README.md +1 -1
  12. package/src/catalog.js +9 -1
  13. package/src/detect.js +28 -9
  14. package/src/install.js +206 -5
  15. package/templates/README.md +1 -1
  16. package/templates/agents/README.md +3 -1
  17. package/templates/agents/agy/README.md +2 -2
  18. package/templates/agents/agy/done-verifier.md +35 -0
  19. package/templates/agents/agy/finding-verifier.md +6 -0
  20. package/templates/agents/agy/reader.md +22 -0
  21. package/templates/agents/claude-code/README.md +7 -5
  22. package/templates/agents/claude-code/builder.md +6 -1
  23. package/templates/agents/claude-code/code-reviewer.md +9 -2
  24. package/templates/agents/claude-code/done-verifier.md +44 -0
  25. package/templates/agents/claude-code/finding-verifier.md +9 -2
  26. package/templates/agents/claude-code/reader.md +26 -0
  27. package/templates/agents/snippets/claude-code.md +9 -4
  28. package/templates/agents/snippets/route-gate.mjs +151 -0
  29. package/templates/agents/snippets/route-metrics.mjs +356 -0
  30. package/templates/agents/snippets/settings.hooks.snippet.json +70 -0
  31. package/templates/agents/snippets/subagent-context.mjs +76 -0
  32. package/templates/beginner/ORCHESTRATOR.md +4 -3
  33. package/templates/common/TASK_BUNDLE.md +2 -2
  34. package/templates/common/protocols/build-protocol.md +2 -2
  35. package/templates/intermediate/ROUTING.md +9 -10
  36. package/templates/intermediate/TIERS.md +2 -0
@@ -1,6 +1,6 @@
1
1
  # Task Bundle: the brief every delegation carries
2
2
 
3
- A subagent, a second CLI, or a fresh chat window holds none of the rules your main session is holding. It cannot see your conventions, it cannot route, and it will read an unspecified edge as an open one.
3
+ A subagent, a second CLI, or a fresh chat window may hold none of the rules your main session is holding, and that is the default to assume. One exception: a Claude Code subagent loads the project's CLAUDE.md hierarchy at start, so it already carries the standing rules, just not this task's scope. Either way, it cannot see this task's conventions and will read an unspecified edge as an open one.
4
4
 
5
5
  > A delegate gets an approved, bounded brief. Absence is not permission.
6
6
 
@@ -25,7 +25,7 @@ Copy this into the delegate's prompt. Delete nothing; write `none` where a field
25
25
  **Denied actions.** <explicit list: do not commit, push, deploy, delete, send, publish, close a ticket...>
26
26
  - Anything absent from Capabilities is denied. Absence is not permission.
27
27
 
28
- **Conventions you do not have.** <restate every house rule this task needs; the delegate holds none>
28
+ **Conventions you do not have.** <restate every house rule this task needs; even a delegate that loaded the standing rules still needs this task's scope, and a second CLI or a fresh chat window may hold none of it>
29
29
 
30
30
  **Report contract.** Return: <exactly what to hand back>. State plainly what you did NOT do
31
31
  and anything you could not verify. "Unverified" is an acceptable answer; a confident guess is not.
@@ -104,14 +104,14 @@ Use a different model family from the one that produced the finding where you ha
104
104
 
105
105
  | Role | Does | Does not |
106
106
  |---|---|---|
107
- | Builder / orchestrator | Routes, maps, writes, verifies, records. Stages 0, 1, 3, 6, 7 | Hand off the main build |
107
+ {{ROLES_BUILDER_ROW}}
108
108
  | Judgment tier | Stage 2 and the architectural arm of Stage 5. Argues with a finished map | Perform the retrieval |
109
109
  | Adversarial auditor | The security arm of Stage 5. Attacks the diff | Fix anything |
110
110
  | Mechanical gates | Stage 4 and any always-on guard | Be overridden without reading |
111
111
  | Cheap workers | Bounded sub-parts: bulk passes, wide searches, long loops | Own a stage |
112
112
  | Human | Stage 5b, and any irreversible or architectural call | Be the first line of review |
113
113
 
114
- **Why the builder does not hand off the main build:** a delegated agent does not inherit the session's standing rules and usually cannot delegate further. Any brief must restate every convention it needs (see `TASK_BUNDLE.md`), and that cost is itself a reason to build directly when the work fits.
114
+ {{BUILDER_HANDOFF_NOTE}}
115
115
 
116
116
  ## Checklist
117
117
 
@@ -15,19 +15,15 @@ Rule of thumb: never spend a frontier token on a task a cheap tier finishes corr
15
15
  0. **Is there a cheaper or better external lane for this?** Check `DELEGATION_MATRIX.md`. Your enabled lanes, every one called through `bin/cli-run.mjs`:
16
16
  {{LANE_STEP0}}
17
17
  1. **Bulk and mechanical?** → fast tier{{BULK_LANE}}. Many independent items each needing its own agent turn → a concurrent fan-out lane if you have one.
18
+ 1a. **Reading or digesting many files or notes, not writing?** → reader. Different from a bulk pass: reader reports, it does not classify, tag or transform.
18
19
  2. **Needs live data?** → {{LIVE_LANE}} standard tier with web tools.
19
20
  3. **Reviewing without changing?** → standard tier read-only. Security-critical → {{ATTACK_LANE}}.
20
21
  3a. **Holding findings from a review or a scanner?** → finding-verifier before any of them cause a repair. A finding is a claim, not a fact.
22
+ 3b. **Checking a tracker item or task against its stated done-signal?** → done-verifier. It probes the named artifact and returns MET, NOT_MET or UNVERIFIABLE; it never closes anything itself.
21
23
  4. **Ambiguous, strategic, expensive to get wrong?** → deep tier (deep-planner). Then hand the plan down.
22
- 5. **Everything else that changes files** → the orchestrator builds it directly. Bounded sub-parts go to cheaper tiers; the main build is never handed off whole.
24
+ {{DECISION_RULE5}}
23
25
 
24
- ## Who builds
25
-
26
- **The orchestrator owns the main build.** It is the only surface that holds these rules: a subagent or a second CLI starts with none of them and cannot route. Handing the main build to one hands it to something the router cannot reach.
27
-
28
- Delegate: background and long-running tasks, small tasks, scoping, verification, research, bounded sub-parts. Never delegate: the main build, or any step that must carry a house rule (secrets handling, the loud-negative verification, the durable record).
29
-
30
- Every delegation carries `TASK_BUNDLE.md`. Its brief must restate every convention the delegate needs.
26
+ {{WHO_BUILDS}}
31
27
 
32
28
  ## The Build Protocol, with lanes bound
33
29
 
@@ -56,7 +52,7 @@ One writer per run; every other lane proposes. Search before writing, index in t
56
52
 
57
53
  ## Modifier rules
58
54
 
59
- - **Plan big, execute small**, within a build: deep tier plans at Checkpoint 1, the orchestrator executes, bulk and wide searches go down.
55
+ {{PLAN_BIG_LINE}}{{INLINE_THRESHOLD_NOTE}}
60
56
  - **Escalation:** never silently retry at the same tier. Escalate one tier or consult deep once, and say which. Two consults that do not unstick it → stop and tell the human.
61
57
  - **De-escalation:** a request that sounds deep but is a lookup routes down.
62
58
  - **Long context:** mechanical digestion → fast tier in chunks; judgment over a long input → standard tier.
@@ -71,8 +67,11 @@ One writer per run; every other lane proposes. Search before writing, index in t
71
67
  |---|---|
72
68
  | "Design the architecture for X" | deep-planner |
73
69
  | "Review this service for bugs" | code-reviewer |
74
- | "Add an endpoint" | the orchestrator builds it |
70
+ {{ADD_ENDPOINT_ROW}}
75
71
  | "Why does this silently drop rows sometimes" | deep-planner (unknown cause), then build the fix directly |
76
72
  | "Summarize these 30 notes into one index" | bulk-worker |
73
+ | "Read every note in this folder and pull out every mention of X" | reader |
77
74
  | "The audit returned 6 findings" | finding-verifier first; repair only what comes back CONFIRMED |
75
+ | "Is issue #123 actually done" | done-verifier |
78
76
  {{LANE_EXAMPLES}}
77
+ {{ROUTE_GATE_SECTION}}
@@ -23,6 +23,8 @@ Tier sets the price per token. Token discipline sets how many tokens. **Effort s
23
23
  | builder | standard | high | a botched deploy is the costly failure |
24
24
  | live-researcher | standard | medium | tools do the retrieval |
25
25
  | bulk-worker | fast | low | the biggest cost win |
26
+ | done-verifier | fast | low | a done-signal check is a lookup, not a judgment call |
27
+ | reader | fast | low | digestion, not judgment |
26
28
 
27
29
  ## Three inputs, not one
28
30