llm-orchestrator 1.3.0 → 1.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (42) hide show
  1. package/.claude-plugin/plugin.json +1 -1
  2. package/README.md +22 -13
  3. package/SKILL.md +3 -1
  4. package/adapters/agents.mjs +111 -12
  5. package/adapters/commands.mjs +2 -2
  6. package/agents/adversarial-skeptic.md +43 -3
  7. package/agents/backend-fixer.md +41 -3
  8. package/agents/code-reviewer.md +43 -3
  9. package/agents/code-simplifier.md +41 -3
  10. package/agents/db-concurrency-specialist.md +45 -3
  11. package/agents/db-migration-author.md +44 -3
  12. package/agents/explore.md +43 -3
  13. package/agents/frontend-fixer.md +41 -3
  14. package/agents/frontend-specialist.md +42 -3
  15. package/agents/general.md +40 -3
  16. package/agents/orchestrator.md +46 -3
  17. package/agents/production-telemetry-collector.md +45 -3
  18. package/agents/provider-webhook-specialist.md +44 -3
  19. package/agents/route-data-flow-tracer.md +43 -3
  20. package/agents/test-engineer.md +42 -3
  21. package/bin/attribution-check.mjs +1 -1
  22. package/commands/incident-close.md +10 -0
  23. package/commands/incident-evidence.md +10 -0
  24. package/commands/incident-fix.md +10 -0
  25. package/commands/incident-start.md +10 -0
  26. package/commands/incident-verify.md +10 -0
  27. package/commands/orchestrate.md +10 -0
  28. package/commands/task-cancel.md +10 -0
  29. package/commands/task-plan.md +10 -0
  30. package/commands/task-status.md +10 -0
  31. package/commands/task-verify.md +10 -0
  32. package/commands/task.md +62 -0
  33. package/lib/adapter-renderer.mjs +5 -1
  34. package/lib/router.mjs +9 -5
  35. package/models/top-models.json +6 -2
  36. package/package.json +2 -1
  37. package/policies/routing.md +2 -2
  38. package/protocol.md +3 -1
  39. package/registries/agent-roles.json +324 -10
  40. package/registries/routing-matrix.json +6 -6
  41. package/schemas/agent-roles.schema.json +88 -1
  42. package/schemas/top-models.schema.json +5 -0
@@ -0,0 +1,62 @@
1
+ ---
2
+ description: "Handle a task through the portable orchestration entrypoint"
3
+ ---
4
+ <!-- llm-orchestrator · created by Bogdan-Gabriel Torcescu · https://www.linkedin.com/in/bogdantorcescu/ · keep this credit when copying or deriving -->
5
+
6
+ Load and follow the `orchestrate-core` skill (or, in a project install, the bridge `.agents/skills/orchestrate/SKILL.md`) before handling this request.
7
+ Mode: execute
8
+ Arguments: $ARGUMENTS
9
+
10
+ Execute the requested task within the project contract.
11
+
12
+ Mandatory — before any planning or execution step, declare the status of each
13
+ core tool in this order (present / gap declared / explicit user refusal):
14
+
15
+ 1. `using-superpowers` (orchestration.bootstrap — invoke first, always)
16
+ 2. `mempalace` recall (memory.recall)
17
+ 3. `mempalace` checkpoint (memory.checkpoint)
18
+ 4. `sequentialthinking` MCP (reasoning.checkpoints)
19
+ 5. `caveman` (communication.concise)
20
+ 6. `rtk` preflight `command -v rtk && rtk --version` (shell.rtk)
21
+ 7. `context7` before code/fix/update/review checks (docs.current)
22
+ 8. `exa` for any external or current fact (research.retrieve)
23
+ 9. `verification-before-completion` (verification.checks)
24
+
25
+ Questions: native, batched, once — every question to the user (mandatory-tool gap, batched
26
+ "skip or fix?" minor findings, destructive confirmation, surviving scope ambiguity) goes through the
27
+ harness's native question mechanism (user.native_question), batched into ONE question at the decision
28
+ point. Never free text at the end of a message; no answer means `blocked_pending_user`, never
29
+ implied approval. Subagents never ask the user: they return `question_for_user` in the handoff.
30
+
31
+ A mandatory gap is declared first and a one-time install is recommended before
32
+ continuing. Only after an explicit user refusal may work continue, and only in
33
+ declared degraded mode (`degraded:` line in every plan, handoff and report) — never silently.
34
+
35
+ Orchestrator: discover tools/permissions (installed / loaded / callable / denied); read the
36
+ project's `## Orchestration bindings (project)` section; allocate task_id; record task/flow
37
+ drawers; emit the pre-evaluation JSON; split the plan into small PlanShards before dispatch
38
+ (one concern, one role, one R/W boundary, one measurable output); dispatch independent shards
39
+ in parallel (2+ agents for MODERATE, 3+ COMPLEX, 4+ CRITICAL when independent work exists;
40
+ parallel groups 2–6, max 8 active shards before synthesis); keep each shard at 8–15 iterations
41
+ and requeue the unfinished remainder at 70%; pass compact drawer refs, scope, owner,
42
+ dependencies, acceptance checks and output drawer — never transcripts; require the
43
+ non-mutating RTK access preflight before each shard; route EVERY shard separately at dispatch
44
+ time against the live inventory, filling its `routing` block (`pair`, `tier`,
45
+ `thinking_level`, `model_requested`, `effort_requested`, `model_effective`,
46
+ `effort_effective`, `review_floor`, `independent_review`, `selection_reason`,
47
+ `inventory_revision`, `price_source`, `est_usd_per_task`) from the W/S/X/F tier and T0–T5
48
+ thinking level in the routing policy, with risk floors and an independent reviewer seat that is
49
+ never cut — never one pair for the whole task, never a silent downgrade below a tier floor
50
+ (`blocked: no eligible model` instead), and on an inventory change re-route the remaining shards
51
+ only; enforce G0–G6; checkpoint to memory every 6–8 tool calls or before
52
+ compaction. If a child returns `Streaming response failed` with a resumable id, record
53
+ `subagent_stream_recovery_pending`, resume that exact child with its original phase contract
54
+ (max three attempts), else one fresh child only from a complete handoff, else
55
+ `subagent_resume_unavailable`. On a required access deny, record
56
+ `permission_recovery:{task_id}:{phase}:{role}`, stop the blocked session and launch one fresh
57
+ session with the matching declared profile (RO/RW); a repeated deny is terminal
58
+ `permission_blocked` — only a human can approve a non-restricted session. Never bypass a
59
+ deny, weaken access grants or skip verification. Integrate only after diff/test evidence; record
60
+ `used_mcps` and verification; run the post-integration cleanup gate before claiming done.
61
+ Sequentialthinking schema: `revisesThought` and `branchFromThought` are integers >= 1
62
+ (use 1 as sentinel when false), never 0/null/omitted.
@@ -28,7 +28,11 @@ Ask the user only through this harness's native question mechanism, batched into
28
28
  Use the \`orchestrate\` skill with the intent \`task\`, \`plan\`, \`status\`, \`cancel\`, or \`verify\` when the harness does not provide a matching native command.
29
29
  `;
30
30
 
31
+ // OpenCode and Kilo load agent files from both `agent/` and `agents/`; the singular form is kept.
31
32
  const AGENT_DIRECTORIES = {claude: '.claude/agents', opencode: '.opencode/agent', kilo: '.kilo/agent'};
33
+ // Frontmatter format per harness; Codex has no native agent-file format, so its
34
+ // generic `.agents/agents` copy uses the Claude format the plugin also ships.
35
+ const AGENT_FORMATS = {claude: 'claude', opencode: 'opencode', kilo: 'kilo', codex: 'claude'};
32
36
 
33
37
  /** Where each harness keeps the flow-adherence hooks: merged JSON, or a plugin file of our own. */
34
38
  export const FLOW_HOOK_TARGETS = {
@@ -112,7 +116,7 @@ export function renderAdapter({harness, capabilities = [], installMode = 'extern
112
116
 
113
117
  if (withAgents) {
114
118
  const agentDirectory = AGENT_DIRECTORIES[harness] ?? '.agents/agents';
115
- for (const agent of agentFiles(agentDirectory)) {
119
+ for (const agent of agentFiles(agentDirectory, AGENT_FORMATS[harness])) {
116
120
  const native = generatedFile({...agent, kind: 'agent-file', existingFiles, ownedPaths: owned});
117
121
  if (native.action === 'conflict') conflicts.push(native.path);
118
122
  else files.push(native);
package/lib/router.mjs CHANGED
@@ -209,10 +209,14 @@ function effortScale(model) {
209
209
  return model.thinking?.control === 'budget_tokens' ? BUDGET_ORDER : EFFORT_ORDER;
210
210
  }
211
211
 
212
- function desiredEffort(model, level) {
212
+ function desiredEffort(model, level, tier = null) {
213
213
  const matrix = loadMatrix();
214
214
  const row = matrix.thinking_levels[level];
215
215
  if (!row) return null;
216
+ // A model seated on a tier below its home runs a declared, lower effort there so it
217
+ // is score-matched to that tier's model rather than over-provisioned.
218
+ const override = tier ? model.thinking?.tier_effort?.[tier]?.[level] : undefined;
219
+ if (override) return override;
216
220
  if (model.thinking?.control === 'budget_tokens') {
217
221
  if (['T0', 'T1'].includes(level)) return 'disabled';
218
222
  if (['T2', 'T3'].includes(level)) return 'enabled';
@@ -228,8 +232,8 @@ function desiredEffort(model, level) {
228
232
  * about the evidence, not licence to invent an enum: we take the model's cheapest
229
233
  * published setting that is at least as deep as the request, or nothing.
230
234
  */
231
- function resolveEffort(model, level) {
232
- const desired = desiredEffort(model, level);
235
+ function resolveEffort(model, level, tier = null) {
236
+ const desired = desiredEffort(model, level, tier);
233
237
  if (desired === null) return { effort: null, substituted: false };
234
238
  const levels = Array.isArray(model.thinking?.levels) ? model.thinking.levels : [];
235
239
  if (levels.length === 0 || levels.includes(desired)) return { effort: desired, substituted: false };
@@ -341,10 +345,10 @@ export function rankModels({
341
345
  if (model.caps?.requires_explicit_flag && !explicitFable51) continue;
342
346
  if (model.caps?.requires_explicit_flag) cap_notes.push(`capped exception: ≤${Math.round((model.caps.share_max ?? 0) * 100)}% of dispatches, explicit request only`);
343
347
 
344
- const resolved = resolveEffort(model, parsed.level);
348
+ const resolved = resolveEffort(model, parsed.level, parsed.tier);
345
349
  let effort = resolved.effort;
346
350
  if (effort === null) continue; // no expression for this thinking level on this control
347
- if (resolved.substituted) cap_notes.push(`no measured \`${desiredEffort(model, parsed.level)}\` point; nearest published setting is \`${effort}\``);
351
+ if (resolved.substituted) cap_notes.push(`no measured \`${desiredEffort(model, parsed.level, parsed.tier)}\` point; nearest published setting is \`${effort}\``);
348
352
 
349
353
  const scale = effortScale(model);
350
354
  const cap = model.caps?.max_effort ?? scale[scale.length - 1];
@@ -366,6 +366,7 @@
366
366
  "admission": "incumbent",
367
367
  "tier": "X",
368
368
  "eligible_tiers": [
369
+ "S",
369
370
  "X"
370
371
  ],
371
372
  "supersedes": "gpt-5-6-sol",
@@ -388,7 +389,10 @@
388
389
  ],
389
390
  "always_on": false,
390
391
  "default_for_tier": "high",
391
- "note": "Same Codex policy as GPT-5.6 Sol: ceiling `high`; T4 is `high` plus an independent second Sol `high` reviewer with no shared history. The API also accepts `none`, which is not measured and never routed."
392
+ "note": "Same Codex policy as GPT-5.6 Sol: ceiling `high`; T4 is `high` plus an independent second Sol `high` reviewer with no shared history. The API also accepts `none`, which is not measured and never routed. On the S seat it runs one notch down (tier_effort), score-matched to Terra.",
393
+ "tier_effort": {
394
+ "S": {"T1": "low", "T2": "low", "T3": "medium"}
395
+ }
392
396
  },
393
397
  "price": {
394
398
  "input_usd_per_mtok": 2,
@@ -463,7 +467,7 @@
463
467
  },
464
468
  "notes": [
465
469
  "Successor to GPT-5.6 Sol on the Codex X seat at half the per-token price ($2/$10 vs $4/$20, permanent pricing). It is cheaper per completed task at every measured effort and scores equal or higher at each (high 43 @ $0.37 vs 42 @ $0.81); GPT-5.6 Sol stays routable as its fallback.",
466
- "Its `medium` (40 @ $0.25) also dominates GPT-5.6 Terra `xhigh` (38 @ $0.63); the S seat stays Terra because an incumbent's tier is its ladder seat, not a score — see the routing note."
470
+ "Also holds the Codex S seat: at every S pair it is cheaper and stronger than GPT-5.6 Terra (S T2: `low` 34 @ $0.13 vs Terra `medium` 30 @ $0.18; S T3: `medium` 40 @ $0.25 vs Terra `high` 34 @ $0.34). Terra stays routable as the S fallback."
467
471
  ],
468
472
  "sources": [
469
473
  "https://artificialanalysis.ai/models/gpt-6-sol",
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "llm-orchestrator",
3
- "version": "1.3.0",
3
+ "version": "1.4.0",
4
4
  "description": "Write /task once — it plans the work, shards it across parallel subagents, gates every phase and verifies before claiming done. Claude Code, Codex, OpenCode, Kilo.",
5
5
  "type": "module",
6
6
  "engines": {
@@ -49,6 +49,7 @@
49
49
  "skills",
50
50
  "hooks",
51
51
  "agents",
52
+ "commands",
52
53
  ".claude-plugin"
53
54
  ],
54
55
  "scripts": {
@@ -75,7 +75,7 @@ the incumbents below are role mappings, not guarantees of availability.
75
75
  | Tier | Role | Claude incumbent | Codex incumbent |
76
76
  | --- | --- | --- | --- |
77
77
  | **W** worker | local, mechanical, repetitive, well-defined: code search, classification, extraction, small edits, boilerplate, simple tests, consistency checks, scoped transforms | Haiku 4.5 (`claude-haiku-4-5`) | `gpt-6-luna`, falling back to `gpt-5.6-luna` |
78
- | **S** standard | default software-engineering model: normal implementation, frontend/backend, moderate debugging, tests, reasonable multi-file refactors, codebase analysis, tool use | Sonnet 5 (`claude-sonnet-5`) | `gpt-5.6-terra` |
78
+ | **S** standard | default software-engineering model: normal implementation, frontend/backend, moderate debugging, tests, reasonable multi-file refactors, codebase analysis, tool use | Sonnet 5 (`claude-sonnet-5`) | `gpt-6-sol` one notch down, falling back to `gpt-5.6-terra` |
79
79
  | **X** senior | hard debugging, architecture, concurrency, migrations, security, auth, payments, billing, backwards compatibility, critical code review, many invariants | Opus 5.5 (`claude-opus-5-5`), falling back to Opus 5 (`claude-opus-5`) | `gpt-6-sol`, falling back to `gpt-5.6-sol` |
80
80
  | **F** frontier | exceptional escalation: very ambiguous, long-horizon, cross-system, major architecture, very large codebase, planning under heavy constraints, or when X fails to produce a solid solution | Fable 5 (`claude-fable-5`) — default F. Fable 5.1 (`claude-fable-5-1`) is a hard-capped exception: **≤2% of all dispatches**, explicit request or documented F-T4 failure on Fable 5 only | GPT-6 Astra (`gpt-6-astra`) — a real single-agent frontier tier, no decomposition workaround needed |
81
81
 
@@ -86,7 +86,7 @@ Verified prices (Sep 2026, provider pricing pages), USD in/out per MTok:
86
86
  | Tier | Claude model | Claude $ in/out | Codex model | Codex $ in/out |
87
87
  | --- | --- | --- | --- | --- |
88
88
  | W | Haiku 4.5 (200K ctx) | $1 / $5 | `gpt-6-luna` (1.05M ctx); fallback `gpt-5.6-luna` $0.20 / $1.20 | $0.10 / $0.50 |
89
- | S | Sonnet 5 (1M ctx) | $2 / $10 | `gpt-5.6-terra` (1.05M ctx) | $2 / $12 |
89
+ | S | Sonnet 5 (1M ctx) | $2 / $10 | `gpt-6-sol` one effort notch down (T1/T2 `low`, T3 `medium`); fallback `gpt-5.6-terra` $2 / $12 | $2 / $10 |
90
90
  | X | Opus 5.5 (1M ctx); fallback Opus 5 $5 / $25 | $4 / $20 | `gpt-6-sol` (1.05M ctx); fallback `gpt-5.6-sol` $4 / $20 | $2 / $10 |
91
91
  | F | Fable 5 (1M ctx) | $10 / $50 | `gpt-6-astra` (1.05M ctx) | $10 / $50 list — **but the cheapest F per completed task of any model here** |
92
92
  | F+ (≤2%) | Fable 5.1 (1M ctx) | $10 / $50 list — **effective cost significantly higher** (always-on thinking, longer turns, more output tokens per task) | — (Astra covers F) | — |
package/protocol.md CHANGED
@@ -46,7 +46,9 @@ flow, but only by declaring it: `llm-orchestrator run start --trivial "<reason>"
46
46
  declaration and its reason are recorded; an undeclared skip is recorded as a skipped flow. A bug
47
47
  fix that needs a regression test, or any change across two or more files, is **not** trivial — it
48
48
  is a typed run. A trivial run that grows past that line gets one reminder to reopen it as a typed
49
- run, and the audit counts it as `trivial_overreach`. A trivial run lasts one turn.
49
+ run, and the audit counts it as `trivial_overreach`. A trivial run lasts one turn. Skipping the
50
+ flow never lifts edit denial: under the orchestrator agent the trivial change is dispatched to
51
+ `general`; without an orchestrator agent it is made by the harness's default build agent.
50
52
 
51
53
  ```json
52
54
  {