llm-orchestrator 1.3.0 → 1.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/README.md +22 -13
- package/SKILL.md +3 -1
- package/adapters/agents.mjs +111 -12
- package/adapters/commands.mjs +2 -2
- package/agents/adversarial-skeptic.md +43 -3
- package/agents/backend-fixer.md +41 -3
- package/agents/code-reviewer.md +43 -3
- package/agents/code-simplifier.md +41 -3
- package/agents/db-concurrency-specialist.md +45 -3
- package/agents/db-migration-author.md +44 -3
- package/agents/explore.md +43 -3
- package/agents/frontend-fixer.md +41 -3
- package/agents/frontend-specialist.md +42 -3
- package/agents/general.md +40 -3
- package/agents/orchestrator.md +46 -3
- package/agents/production-telemetry-collector.md +45 -3
- package/agents/provider-webhook-specialist.md +44 -3
- package/agents/route-data-flow-tracer.md +43 -3
- package/agents/test-engineer.md +42 -3
- package/bin/attribution-check.mjs +1 -1
- package/commands/incident-close.md +10 -0
- package/commands/incident-evidence.md +10 -0
- package/commands/incident-fix.md +10 -0
- package/commands/incident-start.md +10 -0
- package/commands/incident-verify.md +10 -0
- package/commands/orchestrate.md +10 -0
- package/commands/task-cancel.md +10 -0
- package/commands/task-plan.md +10 -0
- package/commands/task-status.md +10 -0
- package/commands/task-verify.md +10 -0
- package/commands/task.md +62 -0
- package/lib/adapter-renderer.mjs +5 -1
- package/lib/router.mjs +9 -5
- package/models/top-models.json +6 -2
- package/package.json +2 -1
- package/policies/routing.md +2 -2
- package/protocol.md +3 -1
- package/registries/agent-roles.json +324 -10
- package/registries/routing-matrix.json +6 -6
- package/schemas/agent-roles.schema.json +88 -1
- package/schemas/top-models.schema.json +5 -0
package/commands/task.md
ADDED
|
@@ -0,0 +1,62 @@
|
|
|
1
|
+
---
|
|
2
|
+
description: "Handle a task through the portable orchestration entrypoint"
|
|
3
|
+
---
|
|
4
|
+
<!-- llm-orchestrator · created by Bogdan-Gabriel Torcescu · https://www.linkedin.com/in/bogdantorcescu/ · keep this credit when copying or deriving -->
|
|
5
|
+
|
|
6
|
+
Load and follow the `orchestrate-core` skill (or, in a project install, the bridge `.agents/skills/orchestrate/SKILL.md`) before handling this request.
|
|
7
|
+
Mode: execute
|
|
8
|
+
Arguments: $ARGUMENTS
|
|
9
|
+
|
|
10
|
+
Execute the requested task within the project contract.
|
|
11
|
+
|
|
12
|
+
Mandatory — before any planning or execution step, declare the status of each
|
|
13
|
+
core tool in this order (present / gap declared / explicit user refusal):
|
|
14
|
+
|
|
15
|
+
1. `using-superpowers` (orchestration.bootstrap — invoke first, always)
|
|
16
|
+
2. `mempalace` recall (memory.recall)
|
|
17
|
+
3. `mempalace` checkpoint (memory.checkpoint)
|
|
18
|
+
4. `sequentialthinking` MCP (reasoning.checkpoints)
|
|
19
|
+
5. `caveman` (communication.concise)
|
|
20
|
+
6. `rtk` preflight `command -v rtk && rtk --version` (shell.rtk)
|
|
21
|
+
7. `context7` before code/fix/update/review checks (docs.current)
|
|
22
|
+
8. `exa` for any external or current fact (research.retrieve)
|
|
23
|
+
9. `verification-before-completion` (verification.checks)
|
|
24
|
+
|
|
25
|
+
Questions: native, batched, once — every question to the user (mandatory-tool gap, batched
|
|
26
|
+
"skip or fix?" minor findings, destructive confirmation, surviving scope ambiguity) goes through the
|
|
27
|
+
harness's native question mechanism (user.native_question), batched into ONE question at the decision
|
|
28
|
+
point. Never free text at the end of a message; no answer means `blocked_pending_user`, never
|
|
29
|
+
implied approval. Subagents never ask the user: they return `question_for_user` in the handoff.
|
|
30
|
+
|
|
31
|
+
A mandatory gap is declared first and a one-time install is recommended before
|
|
32
|
+
continuing. Only after an explicit user refusal may work continue, and only in
|
|
33
|
+
declared degraded mode (`degraded:` line in every plan, handoff and report) — never silently.
|
|
34
|
+
|
|
35
|
+
Orchestrator: discover tools/permissions (installed / loaded / callable / denied); read the
|
|
36
|
+
project's `## Orchestration bindings (project)` section; allocate task_id; record task/flow
|
|
37
|
+
drawers; emit the pre-evaluation JSON; split the plan into small PlanShards before dispatch
|
|
38
|
+
(one concern, one role, one R/W boundary, one measurable output); dispatch independent shards
|
|
39
|
+
in parallel (2+ agents for MODERATE, 3+ COMPLEX, 4+ CRITICAL when independent work exists;
|
|
40
|
+
parallel groups 2–6, max 8 active shards before synthesis); keep each shard at 8–15 iterations
|
|
41
|
+
and requeue the unfinished remainder at 70%; pass compact drawer refs, scope, owner,
|
|
42
|
+
dependencies, acceptance checks and output drawer — never transcripts; require the
|
|
43
|
+
non-mutating RTK access preflight before each shard; route EVERY shard separately at dispatch
|
|
44
|
+
time against the live inventory, filling its `routing` block (`pair`, `tier`,
|
|
45
|
+
`thinking_level`, `model_requested`, `effort_requested`, `model_effective`,
|
|
46
|
+
`effort_effective`, `review_floor`, `independent_review`, `selection_reason`,
|
|
47
|
+
`inventory_revision`, `price_source`, `est_usd_per_task`) from the W/S/X/F tier and T0–T5
|
|
48
|
+
thinking level in the routing policy, with risk floors and an independent reviewer seat that is
|
|
49
|
+
never cut — never one pair for the whole task, never a silent downgrade below a tier floor
|
|
50
|
+
(`blocked: no eligible model` instead), and on an inventory change re-route the remaining shards
|
|
51
|
+
only; enforce G0–G6; checkpoint to memory every 6–8 tool calls or before
|
|
52
|
+
compaction. If a child returns `Streaming response failed` with a resumable id, record
|
|
53
|
+
`subagent_stream_recovery_pending`, resume that exact child with its original phase contract
|
|
54
|
+
(max three attempts), else one fresh child only from a complete handoff, else
|
|
55
|
+
`subagent_resume_unavailable`. On a required access deny, record
|
|
56
|
+
`permission_recovery:{task_id}:{phase}:{role}`, stop the blocked session and launch one fresh
|
|
57
|
+
session with the matching declared profile (RO/RW); a repeated deny is terminal
|
|
58
|
+
`permission_blocked` — only a human can approve a non-restricted session. Never bypass a
|
|
59
|
+
deny, weaken access grants or skip verification. Integrate only after diff/test evidence; record
|
|
60
|
+
`used_mcps` and verification; run the post-integration cleanup gate before claiming done.
|
|
61
|
+
Sequentialthinking schema: `revisesThought` and `branchFromThought` are integers >= 1
|
|
62
|
+
(use 1 as sentinel when false), never 0/null/omitted.
|
package/lib/adapter-renderer.mjs
CHANGED
|
@@ -28,7 +28,11 @@ Ask the user only through this harness's native question mechanism, batched into
|
|
|
28
28
|
Use the \`orchestrate\` skill with the intent \`task\`, \`plan\`, \`status\`, \`cancel\`, or \`verify\` when the harness does not provide a matching native command.
|
|
29
29
|
`;
|
|
30
30
|
|
|
31
|
+
// OpenCode and Kilo load agent files from both `agent/` and `agents/`; the singular form is kept.
|
|
31
32
|
const AGENT_DIRECTORIES = {claude: '.claude/agents', opencode: '.opencode/agent', kilo: '.kilo/agent'};
|
|
33
|
+
// Frontmatter format per harness; Codex has no native agent-file format, so its
|
|
34
|
+
// generic `.agents/agents` copy uses the Claude format the plugin also ships.
|
|
35
|
+
const AGENT_FORMATS = {claude: 'claude', opencode: 'opencode', kilo: 'kilo', codex: 'claude'};
|
|
32
36
|
|
|
33
37
|
/** Where each harness keeps the flow-adherence hooks: merged JSON, or a plugin file of our own. */
|
|
34
38
|
export const FLOW_HOOK_TARGETS = {
|
|
@@ -112,7 +116,7 @@ export function renderAdapter({harness, capabilities = [], installMode = 'extern
|
|
|
112
116
|
|
|
113
117
|
if (withAgents) {
|
|
114
118
|
const agentDirectory = AGENT_DIRECTORIES[harness] ?? '.agents/agents';
|
|
115
|
-
for (const agent of agentFiles(agentDirectory)) {
|
|
119
|
+
for (const agent of agentFiles(agentDirectory, AGENT_FORMATS[harness])) {
|
|
116
120
|
const native = generatedFile({...agent, kind: 'agent-file', existingFiles, ownedPaths: owned});
|
|
117
121
|
if (native.action === 'conflict') conflicts.push(native.path);
|
|
118
122
|
else files.push(native);
|
package/lib/router.mjs
CHANGED
|
@@ -209,10 +209,14 @@ function effortScale(model) {
|
|
|
209
209
|
return model.thinking?.control === 'budget_tokens' ? BUDGET_ORDER : EFFORT_ORDER;
|
|
210
210
|
}
|
|
211
211
|
|
|
212
|
-
function desiredEffort(model, level) {
|
|
212
|
+
function desiredEffort(model, level, tier = null) {
|
|
213
213
|
const matrix = loadMatrix();
|
|
214
214
|
const row = matrix.thinking_levels[level];
|
|
215
215
|
if (!row) return null;
|
|
216
|
+
// A model seated on a tier below its home runs a declared, lower effort there so it
|
|
217
|
+
// is score-matched to that tier's model rather than over-provisioned.
|
|
218
|
+
const override = tier ? model.thinking?.tier_effort?.[tier]?.[level] : undefined;
|
|
219
|
+
if (override) return override;
|
|
216
220
|
if (model.thinking?.control === 'budget_tokens') {
|
|
217
221
|
if (['T0', 'T1'].includes(level)) return 'disabled';
|
|
218
222
|
if (['T2', 'T3'].includes(level)) return 'enabled';
|
|
@@ -228,8 +232,8 @@ function desiredEffort(model, level) {
|
|
|
228
232
|
* about the evidence, not licence to invent an enum: we take the model's cheapest
|
|
229
233
|
* published setting that is at least as deep as the request, or nothing.
|
|
230
234
|
*/
|
|
231
|
-
function resolveEffort(model, level) {
|
|
232
|
-
const desired = desiredEffort(model, level);
|
|
235
|
+
function resolveEffort(model, level, tier = null) {
|
|
236
|
+
const desired = desiredEffort(model, level, tier);
|
|
233
237
|
if (desired === null) return { effort: null, substituted: false };
|
|
234
238
|
const levels = Array.isArray(model.thinking?.levels) ? model.thinking.levels : [];
|
|
235
239
|
if (levels.length === 0 || levels.includes(desired)) return { effort: desired, substituted: false };
|
|
@@ -341,10 +345,10 @@ export function rankModels({
|
|
|
341
345
|
if (model.caps?.requires_explicit_flag && !explicitFable51) continue;
|
|
342
346
|
if (model.caps?.requires_explicit_flag) cap_notes.push(`capped exception: ≤${Math.round((model.caps.share_max ?? 0) * 100)}% of dispatches, explicit request only`);
|
|
343
347
|
|
|
344
|
-
const resolved = resolveEffort(model, parsed.level);
|
|
348
|
+
const resolved = resolveEffort(model, parsed.level, parsed.tier);
|
|
345
349
|
let effort = resolved.effort;
|
|
346
350
|
if (effort === null) continue; // no expression for this thinking level on this control
|
|
347
|
-
if (resolved.substituted) cap_notes.push(`no measured \`${desiredEffort(model, parsed.level)}\` point; nearest published setting is \`${effort}\``);
|
|
351
|
+
if (resolved.substituted) cap_notes.push(`no measured \`${desiredEffort(model, parsed.level, parsed.tier)}\` point; nearest published setting is \`${effort}\``);
|
|
348
352
|
|
|
349
353
|
const scale = effortScale(model);
|
|
350
354
|
const cap = model.caps?.max_effort ?? scale[scale.length - 1];
|
package/models/top-models.json
CHANGED
|
@@ -366,6 +366,7 @@
|
|
|
366
366
|
"admission": "incumbent",
|
|
367
367
|
"tier": "X",
|
|
368
368
|
"eligible_tiers": [
|
|
369
|
+
"S",
|
|
369
370
|
"X"
|
|
370
371
|
],
|
|
371
372
|
"supersedes": "gpt-5-6-sol",
|
|
@@ -388,7 +389,10 @@
|
|
|
388
389
|
],
|
|
389
390
|
"always_on": false,
|
|
390
391
|
"default_for_tier": "high",
|
|
391
|
-
"note": "Same Codex policy as GPT-5.6 Sol: ceiling `high`; T4 is `high` plus an independent second Sol `high` reviewer with no shared history. The API also accepts `none`, which is not measured and never routed."
|
|
392
|
+
"note": "Same Codex policy as GPT-5.6 Sol: ceiling `high`; T4 is `high` plus an independent second Sol `high` reviewer with no shared history. The API also accepts `none`, which is not measured and never routed. On the S seat it runs one notch down (tier_effort), score-matched to Terra.",
|
|
393
|
+
"tier_effort": {
|
|
394
|
+
"S": {"T1": "low", "T2": "low", "T3": "medium"}
|
|
395
|
+
}
|
|
392
396
|
},
|
|
393
397
|
"price": {
|
|
394
398
|
"input_usd_per_mtok": 2,
|
|
@@ -463,7 +467,7 @@
|
|
|
463
467
|
},
|
|
464
468
|
"notes": [
|
|
465
469
|
"Successor to GPT-5.6 Sol on the Codex X seat at half the per-token price ($2/$10 vs $4/$20, permanent pricing). It is cheaper per completed task at every measured effort and scores equal or higher at each (high 43 @ $0.37 vs 42 @ $0.81); GPT-5.6 Sol stays routable as its fallback.",
|
|
466
|
-
"
|
|
470
|
+
"Also holds the Codex S seat: at every S pair it is cheaper and stronger than GPT-5.6 Terra (S T2: `low` 34 @ $0.13 vs Terra `medium` 30 @ $0.18; S T3: `medium` 40 @ $0.25 vs Terra `high` 34 @ $0.34). Terra stays routable as the S fallback."
|
|
467
471
|
],
|
|
468
472
|
"sources": [
|
|
469
473
|
"https://artificialanalysis.ai/models/gpt-6-sol",
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "llm-orchestrator",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.4.0",
|
|
4
4
|
"description": "Write /task once — it plans the work, shards it across parallel subagents, gates every phase and verifies before claiming done. Claude Code, Codex, OpenCode, Kilo.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"engines": {
|
|
@@ -49,6 +49,7 @@
|
|
|
49
49
|
"skills",
|
|
50
50
|
"hooks",
|
|
51
51
|
"agents",
|
|
52
|
+
"commands",
|
|
52
53
|
".claude-plugin"
|
|
53
54
|
],
|
|
54
55
|
"scripts": {
|
package/policies/routing.md
CHANGED
|
@@ -75,7 +75,7 @@ the incumbents below are role mappings, not guarantees of availability.
|
|
|
75
75
|
| Tier | Role | Claude incumbent | Codex incumbent |
|
|
76
76
|
| --- | --- | --- | --- |
|
|
77
77
|
| **W** worker | local, mechanical, repetitive, well-defined: code search, classification, extraction, small edits, boilerplate, simple tests, consistency checks, scoped transforms | Haiku 4.5 (`claude-haiku-4-5`) | `gpt-6-luna`, falling back to `gpt-5.6-luna` |
|
|
78
|
-
| **S** standard | default software-engineering model: normal implementation, frontend/backend, moderate debugging, tests, reasonable multi-file refactors, codebase analysis, tool use | Sonnet 5 (`claude-sonnet-5`) | `gpt-5.6-terra` |
|
|
78
|
+
| **S** standard | default software-engineering model: normal implementation, frontend/backend, moderate debugging, tests, reasonable multi-file refactors, codebase analysis, tool use | Sonnet 5 (`claude-sonnet-5`) | `gpt-6-sol` one notch down, falling back to `gpt-5.6-terra` |
|
|
79
79
|
| **X** senior | hard debugging, architecture, concurrency, migrations, security, auth, payments, billing, backwards compatibility, critical code review, many invariants | Opus 5.5 (`claude-opus-5-5`), falling back to Opus 5 (`claude-opus-5`) | `gpt-6-sol`, falling back to `gpt-5.6-sol` |
|
|
80
80
|
| **F** frontier | exceptional escalation: very ambiguous, long-horizon, cross-system, major architecture, very large codebase, planning under heavy constraints, or when X fails to produce a solid solution | Fable 5 (`claude-fable-5`) — default F. Fable 5.1 (`claude-fable-5-1`) is a hard-capped exception: **≤2% of all dispatches**, explicit request or documented F-T4 failure on Fable 5 only | GPT-6 Astra (`gpt-6-astra`) — a real single-agent frontier tier, no decomposition workaround needed |
|
|
81
81
|
|
|
@@ -86,7 +86,7 @@ Verified prices (Sep 2026, provider pricing pages), USD in/out per MTok:
|
|
|
86
86
|
| Tier | Claude model | Claude $ in/out | Codex model | Codex $ in/out |
|
|
87
87
|
| --- | --- | --- | --- | --- |
|
|
88
88
|
| W | Haiku 4.5 (200K ctx) | $1 / $5 | `gpt-6-luna` (1.05M ctx); fallback `gpt-5.6-luna` $0.20 / $1.20 | $0.10 / $0.50 |
|
|
89
|
-
| S | Sonnet 5 (1M ctx) | $2 / $10 | `gpt-5.6-terra`
|
|
89
|
+
| S | Sonnet 5 (1M ctx) | $2 / $10 | `gpt-6-sol` one effort notch down (T1/T2 `low`, T3 `medium`); fallback `gpt-5.6-terra` $2 / $12 | $2 / $10 |
|
|
90
90
|
| X | Opus 5.5 (1M ctx); fallback Opus 5 $5 / $25 | $4 / $20 | `gpt-6-sol` (1.05M ctx); fallback `gpt-5.6-sol` $4 / $20 | $2 / $10 |
|
|
91
91
|
| F | Fable 5 (1M ctx) | $10 / $50 | `gpt-6-astra` (1.05M ctx) | $10 / $50 list — **but the cheapest F per completed task of any model here** |
|
|
92
92
|
| F+ (≤2%) | Fable 5.1 (1M ctx) | $10 / $50 list — **effective cost significantly higher** (always-on thinking, longer turns, more output tokens per task) | — (Astra covers F) | — |
|
package/protocol.md
CHANGED
|
@@ -46,7 +46,9 @@ flow, but only by declaring it: `llm-orchestrator run start --trivial "<reason>"
|
|
|
46
46
|
declaration and its reason are recorded; an undeclared skip is recorded as a skipped flow. A bug
|
|
47
47
|
fix that needs a regression test, or any change across two or more files, is **not** trivial — it
|
|
48
48
|
is a typed run. A trivial run that grows past that line gets one reminder to reopen it as a typed
|
|
49
|
-
run, and the audit counts it as `trivial_overreach`. A trivial run lasts one turn.
|
|
49
|
+
run, and the audit counts it as `trivial_overreach`. A trivial run lasts one turn. Skipping the
|
|
50
|
+
flow never lifts edit denial: under the orchestrator agent the trivial change is dispatched to
|
|
51
|
+
`general`; without an orchestrator agent it is made by the harness's default build agent.
|
|
50
52
|
|
|
51
53
|
```json
|
|
52
54
|
{
|