@uzysjung/agent-harness 26.152.0 → 26.154.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (28) hide show
  1. package/dist/{chunk-EQDC2AAU.js → chunk-242TVGMZ.js} +582 -556
  2. package/dist/chunk-242TVGMZ.js.map +1 -0
  3. package/dist/index.js +623 -303
  4. package/dist/index.js.map +1 -1
  5. package/dist/trust-tier-drift.js +1 -1
  6. package/package.json +1 -1
  7. package/templates/CLAUDE.md +93 -145
  8. package/templates/codex/README.md +6 -7
  9. package/templates/codex/config.toml.template +1 -9
  10. package/templates/skills/audit-harness-fit/README.md +76 -43
  11. package/templates/skills/audit-harness-fit/SKILL.md +62 -47
  12. package/templates/skills/audit-harness-fit/evals/scenarios.yaml +179 -22
  13. package/templates/skills/audit-harness-fit/references/apply.md +74 -45
  14. package/templates/skills/audit-harness-fit/references/audit.md +142 -118
  15. package/templates/skills/audit-harness-fit/references/populate.md +53 -49
  16. package/templates/skills/audit-harness-fit/references/verification.md +131 -100
  17. package/templates/skills/clear-korean-communication/SKILL.md +76 -209
  18. package/templates/skills/external-model-consult/scripts/codex-ask.sh +9 -4
  19. package/templates/skills/external-model-consult/scripts/gemini-ask.sh +9 -4
  20. package/templates/skills/model-orchestration/SKILL.md +123 -267
  21. package/templates/skills/objective-brief/SKILL.md +3 -3
  22. package/dist/chunk-EQDC2AAU.js.map +0 -1
  23. package/templates/codex/hooks/README.md +0 -37
  24. package/templates/codex/hooks/session-start.sh +0 -7
  25. package/templates/codex/hooks/uncommitted-check.sh +0 -7
  26. package/templates/skills/clear-korean-communication/references/pre-send-checklist.md +0 -37
  27. package/templates/skills/clear-korean-communication/references/why-it-works.md +0 -24
  28. package/templates/skills/clear-korean-communication/references/worked-examples.md +0 -89
@@ -146,17 +146,22 @@ env -i PATH="$PATH" HOME="$HOME" TERM="${TERM:-dumb}" \
146
146
  "$CODEX" "${ARGS[@]}" 1>&2 &
147
147
  CPID=$!
148
148
  TIMED_OUT=0
149
- ELAPSED=0
149
+ # Poll at 200 ms, not 1 s (#411): with a 1 s tick every call paid up to a full second of idle
150
+ # wait after the CLI had already exited, which is what pushed the wrapper's test cases toward
151
+ # vitest's 5 s budget on a busy machine. TICKS counts fifths of a second; both BSD and GNU
152
+ # `sleep` accept the fractional argument.
153
+ TICKS=0
154
+ TICK_LIMIT=$((TIMEOUT_S * 5))
150
155
  while kill -0 "$CPID" 2>/dev/null; do
151
- if [ "$ELAPSED" -ge "$TIMEOUT_S" ]; then
156
+ if [ "$TICKS" -ge "$TICK_LIMIT" ]; then
152
157
  TIMED_OUT=1
153
158
  kill "$CPID" 2>/dev/null || true
154
159
  sleep 2
155
160
  kill -9 "$CPID" 2>/dev/null || true
156
161
  break
157
162
  fi
158
- sleep 1
159
- ELAPSED=$((ELAPSED + 1))
163
+ sleep 0.2
164
+ TICKS=$((TICKS + 1))
160
165
  done
161
166
  set +e
162
167
  wait "$CPID"
@@ -183,17 +183,22 @@ env -i PATH="$PATH" HOME="$HOME" TERM="${TERM:-dumb}" \
183
183
  "$AGY" --model="$MODEL" -p "$PROMPT" >"$MSGFILE" &
184
184
  APID=$!
185
185
  TIMED_OUT=0
186
- ELAPSED=0
186
+ # Poll at 200 ms, not 1 s (#411): with a 1 s tick every call paid up to a full second of idle
187
+ # wait after the CLI had already exited, which is what pushed the wrapper's test cases toward
188
+ # vitest's 5 s budget on a busy machine. TICKS counts fifths of a second; both BSD and GNU
189
+ # `sleep` accept the fractional argument.
190
+ TICKS=0
191
+ TICK_LIMIT=$((TIMEOUT_S * 5))
187
192
  while kill -0 "$APID" 2>/dev/null; do
188
- if [ "$ELAPSED" -ge "$TIMEOUT_S" ]; then
193
+ if [ "$TICKS" -ge "$TICK_LIMIT" ]; then
189
194
  TIMED_OUT=1
190
195
  kill "$APID" 2>/dev/null || true
191
196
  sleep 2
192
197
  kill -9 "$APID" 2>/dev/null || true
193
198
  break
194
199
  fi
195
- sleep 1
196
- ELAPSED=$((ELAPSED + 1))
200
+ sleep 0.2
201
+ TICKS=$((TICKS + 1))
197
202
  done
198
203
  set +e
199
204
  wait "$APID"
@@ -1,273 +1,129 @@
1
1
  ---
2
2
  name: model-orchestration
3
3
  description: >-
4
- Apply the fixed model-role and thinking-effort policy whenever work is delegated to subagents
5
- or a model/effort choice is made: the orchestrator (top-tier model, Fable) DIRECTLY owns
6
- 설계·기획·분배·리뷰 — sets direction, reviews plan/spec docs, improves shipped features, and hunts
7
- perf/security problems; core implementation, test authoring/execution
8
- (E2E included), and code verification/V&V go to Opus at xhigh or above;
9
- repetitive/simple implementation goes to Sonnet at high or above — Sonnet is never used for
10
- tests or verification; plan/spec drafts may be produced by Opus from a Fable direction brief,
11
- but the review/decision is always Fable's. Never delegate below the effort floors. Use whenever
12
- you are about to spawn an Agent/Task/Workflow worker, pick a model for a subtask, set a
13
- thinking/effort level, or assign verification. Trigger on "위임해", "오케스트레이션", "모델 역할분담", "어떤 모델로",
14
- "effort 얼마로", "서브에이전트", "delegate this", "which model should". Fire even when the
15
- user doesn't name the policy — any delegation is in scope.
4
+ Select models, reasoning effort, and delegation boundaries for a task. Use when
5
+ choosing direct execution versus subagents, allocating parallel work or
6
+ independent review, or reducing multi-agent token and context overhead. Favor
7
+ context reuse and add model calls or workers when their expected contribution
8
+ justifies the cost.
16
9
  ---
17
10
 
18
- # Model Orchestration Policy
11
+ # Model Orchestration
19
12
 
20
- **What is policy and what is vocabulary.** The policy is the role split — orchestrator ≠ builder ≠
21
- verifier — and the rule that verification never goes to a lower tier than the work it verifies.
22
- `Fable` · `Opus` · `Sonnet` and the effort levels (`xhigh` · `high` · `max`) are Claude Code
23
- vocabulary. On Codex, OpenCode, or Antigravity apply the same split with that vendor's top tier for
24
- orchestration and verification and its mid tier for repetitive implementation; an effort knob that
25
- does not exist there is not a violation — putting verification on a lower tier is.
26
-
27
- A fixed role split between model tiers, set by the user (2026-07-04, revised 2026-07-07 and
28
- 2026-08-02 — *"설계·분배·기획·리뷰는 Fable5, 핵심 구현·테스트·검증은 opus5, 반복·단순 구현은
29
- sonnet"*). The premise is
30
- **quality-over-cost**: delegation floors are set at the effort levels Anthropic itself
31
- recommends for intelligence-sensitive work ("start with xhigh for coding and agentic use
32
- cases, high as the minimum for most intelligence-sensitive workloads" — official effort
33
- guidance), not at the cost-saving end ("low ... like subagents"). When this policy and a
34
- cost instinct conflict, the policy wins; surface the cost, don't silently downgrade.
35
-
36
- ## Role split
37
-
38
- | Role | Who | Effort floor | Duties |
39
- |------|-----|--------------|--------|
40
- | **Orchestrator + Product Manager** | Top-tier session model (Fable) | session level | 설계·기획·분배·**리뷰** 전부. Directly: 서비스 방향성 수립·논의, 설계 결정·트레이드오프 확정, 작업 분배, 기획/스펙문서 **리뷰·중재**, 기개발 기능 **개선**, 문제점(성능·보안) 발굴 — see below |
41
- | **Core builder + Tester + V&V** | `opus` | **xhigh** (or `max`) | **핵심 구현**(새 설계 표면·경계 판단이 남는 것) + **테스트 작성·실행**(E2E 포함) + **코드 리뷰·검증/V&V**(fresh instance). 기획/스펙 문서 *초안*은 Fable의 방향 브리프를 받아 작성 가능하되, 확정은 Fable 리뷰를 거친다. **리뷰의 분업**: 설계·기획·스펙 문서 리뷰 = Fable / 코드(diff) 리뷰 = Opus V&V |
42
- | **Repetitive builder + read-only assist** | `sonnet` | **high** (or above) | **반복·단순 구현** — 확립된 패턴의 적용, 설계 판단이 이미 끝난 기계적 구현 (산출물은 Opus 교차검증) + 리서치 스윕. **테스트·검증 투입 금지** |
43
- | **Orchestrator stand-in** (quota exhausted) | `opus` @ `max` | max | Takes over orchestration via a handoff — see "Orchestrator handoff" |
44
-
45
- The orchestrator assigns a thinking/effort level **per task** at delegation time. Delegating
46
- below a floor is a policy violation, not a tuning choice.
47
-
48
- ## The orchestrator's own lane: direction, review, judgment
49
-
50
- The orchestrator authors **direction**, not documents. 서비스 방향성(where the product goes),
51
- 스펙 리뷰(is this plan right?), 기능 개선(what should get better in what shipped?), and
52
- 성능·보안 문제 발굴 stay in the orchestrator's window — that's where user intent, constraints,
53
- and history live, and judgment is the one thing a delegation prompt can't carry.
54
-
55
- Document *drafting* is delegated to Opus with a direction brief (intent, constraints, decided
56
- trade-offs). This buys an author≠reviewer split for documents themselves: **Opus writes, the
57
- orchestrator reviews** — a draft you review with fresh eyes gets scrutiny your own draft never
58
- would. For 기획/스펙문서 리뷰, run the [[multi-persona-review]] skill (3–5 disjoint persona
59
- reviewers, severity-ranked synthesis) instead of a single-pass read; the orchestrator arbitrates
60
- its findings rather than line-editing alone.
61
-
62
- ## Routing test
63
-
64
- Before routing to any model: **deterministic transforms don't get a model at all.** A rename,
65
- a format sweep, a mechanical find-and-replace is `sed`/`grep`/script work — spending model
66
- tokens on what code answers deterministically fails the routing test at step zero.
67
-
68
- Past that gate (2026-08-02 사용자 결정):
69
-
70
- - **Core implementation, all tests, anything that judges code** — new design surface, boundary
71
- decisions, test authoring and execution (E2E included), code verification/V&V — route to
72
- **Opus @ xhigh+**.
73
- - **Repetitive/simple implementation** — applying an established pattern where the design
74
- decisions are already made (N번째 유사 변환, 확정 스펙의 기계적 구현) — routes to
75
- **Sonnet @ high+**, and its output gets Opus cross-verification before it gates anything.
76
- The routing question is *"새 판단이 남아 있는가?"* — 남아 있으면 핵심(Opus), 없으면 반복(Sonnet).
77
- - **Read-only assists** (research sweeps) may also go to Sonnet; tests and verification never do.
78
- - **Judgment-free implementation whose pass condition is a command** — the repetitive lane above
79
- stays the default; only when every predicate in the next section holds may that work go to an
80
- external executor instead.
81
- - **Korean the user will read** · **shorter and better structured** — an external advisory
82
- round-trip. Which provider takes which is `external-model-consult`'s provider table; that table
83
- is the SSOT and is not copied here.
84
- - **A judgment that needs several perspectives** → [[multi-persona-review]] (native panel). **One
85
- non-Claude perspective** → `external-model-consult` persona mode. A panel that mixes both asks
86
- the user before it runs, and that gate belongs to [[multi-persona-review]].
87
- - 설계·기획·분배·리뷰 decisions never route down at all — they are the orchestrator's own
88
- (see above).
89
-
90
- Parallelism rule of thumb (three independent sources converge on this): **parallel reads are
91
- safe, parallel writes are dangerous**. Fan out freely for research/search/review; keep writes
92
- sequential or isolated (worktree) so two workers never make conflicting implicit decisions in
93
- the same files.
94
-
95
- ## External executors — the lane outside the harness
96
-
97
- This lane buys **capacity, not quality.** It is not a fourth rank below Sonnet — the ranking above
98
- is by judgment, and an outside CLI has not earned a place in it. What it can hold is work whose
99
- **quality a command decides**, the only case where "cheaper" doesn't also mean "worse." Choosing it
100
- because some other vendor is supposedly better inverts this policy's quality-over-cost premise.
101
-
102
- **All five predicates must hold; any one false closes the lane.** Each is a question you can answer
103
- right now, not a "use it if it helps."
104
-
105
- | # | Predicate | How you answer it now |
106
- |---|---|---|
107
- | **P1** | No judgment is left | The routing question above, unchanged — *"새 판단이 남아 있는가?"* 남아 있으면 닫힌다 |
108
- | **P2** | Passing is machine-decided | Write the pass command on one line, right now. Can't? Closed |
109
- | **P3** | The output gates nothing | It gates no merge, no release, no decision until in-harness cross-verification clears it |
110
- | **P4** | The repo may go to that provider | The first-use approval below is done, and this file set is inside what was approved |
111
- | **P5** | It is **not the CLI you are running on**, and it can use a shell | Delegating to yourself is not a round trip; with no shell, this lane doesn't exist for you |
112
-
113
- Even with all five true, **the in-harness repetitive lane is still the default.** Go outside when
114
- mechanical work is eating capacity a judgment lane needs — and reach for the fewest tools that
115
- answer the task, not every tool installed.
116
-
117
- **First use in a repository is the user's call, not yours.** Routing implementation to an external
118
- CLI puts this repository's code into another vendor's session — a disclosure none of the in-harness
119
- lanes make. Before the first such delegation in a project, say which tool, which provider its own
120
- config resolves to, which files the worker may touch, and what comes back; then wait. After that one
121
- approval, routing inside the predicates above is yours. Ask again when the boundary moves — a
122
- different tool, or files outside what was approved.
123
-
124
- **Tool missing, auth expired, provider refused → step down a lane and report what you could not
125
- use.** Never install it, never log in for them, and never quietly substitute a different provider:
126
- the user knows which tool answered, so a silent swap makes your report false. Recognizing each
127
- failure — and the exact wording for it — belongs to [[external-model-consult]]; where that skill
128
- isn't installed, only the conclusion survives: stop and ask.
129
-
130
- **This lane does not choose models.** The tool runs whatever its own config resolves to, and you
131
- report what answered. A model id written down here goes stale and pins the user to a retired model.
132
-
133
- **Call the tool's non-interactive mode from the shell.** Read the subcommand off `--help` rather
134
- than typing one from memory, and stop and report if it isn't there. Writes stay isolated (worktree
135
- or equivalent) — the parallel-write rule above holds for outside workers too.
136
-
137
- The delegation prompt spec and the file-handoff contract below apply here unchanged: an outside
138
- worker is still a worker.
139
-
140
- ## Effort floors — and the inheritance gotcha
141
-
142
- Effort tiers: `low` < `medium` < `high` < `xhigh` < `max`. Fable 5, Sonnet 5, and Opus 4.8/4.7
143
- support all five; on models without `xhigh` the request silently falls back to the nearest
144
- lower level — check the model before assuming the floor holds.
145
-
146
- **The gotcha that breaks this policy silently:** the Agent/Task tool accepts a per-invocation
147
- `model`, but **not a per-invocation `effort`** — a spawned agent inherits the session effort
148
- unless its definition says otherwise. Session default is `high`, so a plain
149
- `Agent(model: "opus")` runs at `high`, **below the xhigh floor**, with no warning. Enforce the
150
- floor through one of these three paths:
151
-
152
- 1. **Pre-defined agent roles** — `.claude/agents/<role>.md` frontmatter pins both knobs.
153
- This is the durable path for recurring roles:
154
-
155
- ```yaml
156
- ---
157
- name: verifier
158
- description: Fresh-context V&V per model-orchestration policy
159
- model: opus
160
- effort: xhigh
161
- ---
162
- ```
163
-
164
- 2. **Workflow scripts** — `agent(prompt, {model: "opus", effort: "xhigh"})` supports both
165
- per call. Use for scripted fan-outs.
166
-
167
- 3. **Session inheritance** — if the session already runs at `xhigh` (e.g. `/effort xhigh` or
168
- ultracode), a bare `Agent(model: "opus")` inherits a compliant level. Verify with
169
- `/effort`; don't assume.
170
-
171
- One environment caveat: `CLAUDE_CODE_SUBAGENT_MODEL` outranks every per-invocation and
172
- frontmatter model choice. If delegation models look wrong, check that env var first.
173
-
174
- ## Delegation prompt spec
175
-
176
- [[objective-brief]] owns the delegation prompt format — write the brief there and hand it over.
177
- Without that skill installed, the five elements are: objective (and why it matters), output
178
- format, tool/source guidance, boundaries, acceptance criteria.
179
-
180
- Scale worker count to the task, stated up front in your own plan: trivial lookup → no agent at
181
- all (do it directly); bounded question → one agent; genuinely independent axes → one agent per
182
- axis. Over-spawning is a documented failure mode, and multi-agent runs cost ~15× a plain chat
183
- turn — delegate when the task's value justifies it, not by reflex.
184
-
185
- ## Worker lifecycle — 다 쓴 에이전트는 닫는다
186
-
187
- Closing what you spawned is owned by the `git-policy` rule's Session Cleanup section — the
188
- sweep there uses an enumeration probe rather than memory, because compaction wipes the
189
- orchestrator's own ledger of open workers.
190
-
191
- ### Collect results as a file, not as a return message
192
-
193
- A worker's final message is not durable transport — long results get truncated silently, and
194
- what was cut is gone. This burned three separate times — most
195
- recently a panel run that lost three reviewers' output — before it was codified as a gate.
196
- The contract: the worker writes its deliverable to a scratchpad/output **file whose path the
197
- orchestrator fixes in the spawn prompt** — collection is a spawn-time decision, because a
198
- truncated reply cannot be retrofit into a file after the fact. The final message then carries
199
- only the file path plus a short summary. Treat any delegation whose prompt names no output
200
- file as not yet dispatched.
201
-
202
- ## V&V separation
203
-
204
- The implementer never verifies its own work — the *instance* that wrote something never judges
205
- it. How that plays out per lane:
206
-
207
- - **Sonnet implemented** (반복·단순 구현) → **Opus cross-verifies** — a different model has
208
- different blind spots, which is the cheapest anchoring break available.
209
- - **Opus implemented** (핵심 구현·테스트) → a **fresh Opus instance** verifies: a NEW agent with no
210
- shared history. A verifier that watched the implementation happen inherits the implementer's
211
- mental model and anchors on it — same-session self-review reliably misses the same edge
212
- cases the implementation missed. Instance separation is what makes "Opus builds AND Opus
213
- verifies" coherent.
214
- - **Orchestrator layer on top**: the orchestrator hunts 성능·보안 문제점 in shipped features —
215
- a second, higher-altitude pass that judges what a diff-level verifier doesn't (product fit,
216
- systemic risk). Documents get the same split: Opus authors, the orchestrator reviews via
217
- [[multi-persona-review]].
218
-
219
- This pairs with, not replaces, deterministic gates (tests, typecheck, CI) — the verifier
220
- judges what automation can't: spec fit, missed edge cases, design drift.
221
-
222
- ## Orchestrator handoff (quota exhaustion)
223
-
224
- When the top-tier orchestrator's quota runs out mid-project, **Opus @ `max` takes over
225
- orchestration**. There is no reliable automatic path — documented fallback chains explicitly
226
- exclude rate-limit errors, and plan-level auto-switching is undocumented behavior you must not
227
- build on. Hand off manually:
228
-
229
- 1. Run the [[compaction-handoff]] protocol: persist durable facts to memory, take an atomic
230
- git snapshot (clean tree + open-PR check), emit the fixed-field resume anchor
231
- (current state / verified / what's left / next action).
232
- 2. The successor session starts on `opus` at `max` (`/model opus` + `/effort max`), reads the
233
- anchor, and continues as orchestrator under this same policy — setting direction, reviewing,
234
- and hunting problems itself (it can author documents directly too, since it already runs at
235
- the Opus doc-author tier — no separate delegation needed while it stands in).
236
- 3. When the top-tier model becomes available again, hand back the same way.
237
-
238
- ## Anti-patterns
239
-
240
- | Anti-pattern | Why it's a violation |
241
- |---|---|
242
- | `Agent(model: "opus")` with session at default effort | Inherits `high` < xhigh floor — use a pinned agent role, Workflow opts, or raise session effort |
243
- | Delegating 방향성 수립 or a final judgment call to a worker | Direction is the orchestrator's own — a delegation prompt can't carry the shaping context |
244
- | Accepting an Opus-authored spec/plan without orchestrator review | Author≠reviewer applies to documents too — run [[multi-persona-review]] before accepting |
245
- | Orchestrator hand-writing full spec/plan drafts itself | Authoring is Opus's lane — brief the direction, delegate the draft, review the result |
246
- | Sonnet on tests, verification, or **core** implementation | Sonnet 의 레인은 반복·단순 구현(확립 패턴 적용)과 읽기 전용 보조까지다 — 테스트·V&V·새 설계 판단은 Opus/Fable 몫 |
247
- | Sonnet 반복 구현물이 Opus 교차검증 없이 머지를 게이트 | 교차검증이 Sonnet 레인의 존립 조건이다 — 검증 없는 반복 구현은 레인 위반과 같다 |
248
- | Effort below floor "to save tokens" | The floors ARE the policy; surface cost concerns to the user instead |
249
- | Implementer verifying its own diff | Anchoring — verification needs fresh context, prefer a different model |
250
- | Two agents writing the same files in parallel | Conflicting implicit decisions; keep writes sequential or worktree-isolated |
251
- | Spawning an agent for what one direct tool call answers | 15× token multiplier for zero value — do trivial work directly |
252
- | Relying on plan-level auto-fallback for continuity | Undocumented behavior; use the manual handoff protocol |
253
- | Finished worker left running after its result is consumed | Subagent panes/windows accumulate (iTerm2 등) and idle pings pollute the session — TaskStop as one motion with consuming the result |
254
- | 외부 실행기에 **테스트 작성·검증**·핵심 구현을 넘김 | 이 레인은 판단 잔여 0 인 일만 받는다 — 무엇을 단언할지 정하는 일을 밖으로 내보내면 외부 산출물을 검사할 기준 자체가 밖에 있게 된다. 형태가 이미 고정된 표에 케이스 한 줄을 복제하는 일은 무엇을 단언할지 정하지 않으므로 여기 해당하지 않는다 |
255
- | 도구가 없어서 조용히 다른 제공자로 갈아타 실행 | 사용자는 어느 도구가 답했는지 알고 있다 — 대체는 보고 대상이지 판단 대상이 아니다 |
256
- | 외부 실행기를 "품질이 더 낫다"는 이유로 고름 | 이 레인이 사는 것은 용량이다. 품질을 근거로 들면 이 정책의 quality-over-cost 전제를 뒤집는 것이다 |
257
-
258
- ## Quick reference
259
-
260
- ```
261
- 방향성 수립 / 설계·기획 확정 / 작업 분배 / 스펙 리뷰(multi-persona-review) / 기능 개선 / 성능·보안 문제발굴
262
- → 오케스트레이터(Fable) 직접
263
- 스펙 문서 초안(Fable 브리프 기반) / 핵심 구현 / 테스트 작성·실행(E2E 포함) / 코드 리뷰·검증(V&V)
264
- → opus @ xhigh (또는 max) — pinned role 또는 Workflow opts
265
- 반복·단순 구현(확립 패턴 적용) / 리서치 스윕
266
- → sonnet @ high 이상 — 테스트·검증 투입 금지, 산출물은 Opus 교차검증
267
- 결정적 변환 (rename·포맷) → 모델 위임 금지 — sed/grep/스크립트 직접
268
- 판단 잔여 0 + 합격을 명령 하나로 판정
269
- → sonnet 기본 / 다섯 술어 충족 시 외부 실행기 — 최초 1회 사용자 확인, 산출물은 in-harness 교차검증
270
- 도구 부재·인증 만료 → 레인을 내리고 무엇을 못 썼는지 보고 — 대신 설치/로그인 금지, 조용한 제공자 교체 금지
271
- 위임 완료 → 결과 수거와 동시에 TaskStop (SendMessage 재사용 예정 시만 유지 선언)
272
- Fable 소진 → compaction-handoff → opus @ max 가 오케스트레이터 대행
273
- ```
13
+ Apply intelligence where it most improves a verified, usable outcome. Optimize
14
+ whole-task cost and delay, including context duplication, coordination,
15
+ verification, and likely rework, while preserving required quality and authority.
16
+ These are allocation principles, not a mandatory multi-agent workflow. Explicit
17
+ project requirements remain binding.
18
+
19
+ ## 1. Keep ownership with useful context
20
+
21
+ The agent holding the user's intent and relevant project context owns the
22
+ integrated result. It may plan, draft, implement, run checks, and refine directly.
23
+ Roles describe responsibilities; they do not require separate workers.
24
+
25
+ Keep tightly coupled work together when a handoff would mostly make another
26
+ agent reconstruct the same understanding. Completing the task in one capable
27
+ context is a valid orchestration choice, not a failure to delegate.
28
+
29
+ Treat model capability, reasoning effort, and context count as separate choices.
30
+ Where supported, changing model or effort within an existing context may be more
31
+ useful than spawning another agent. Reuse applicable routing decisions rather
32
+ than reopening them for every tool call or phase.
33
+
34
+ ## 2. Put capability at the bottleneck
35
+
36
+ Use stronger reasoning for unresolved design choices, difficult diagnosis,
37
+ cross-boundary behavior, and consequential mistakes. Choose sufficient capability
38
+ early when uncertainty or recovery cost warrants it; there is no need to exhaust
39
+ weaker models first. A stronger model doing the work directly may cost less
40
+ overall than a weaker implementation followed by repeated review and repair.
41
+
42
+ For well-specified work, choose between direct execution, deterministic tools,
43
+ and a lighter capable model by total completion cost. Running a known test command
44
+ and deciding whether its coverage is adequate are different demands on judgment.
45
+ Use tools for execution and measurement; allocate reasoning to decisions they
46
+ cannot settle.
47
+
48
+ Choose from actual capabilities and task evidence rather than fixed vendor-role
49
+ assignments or universal effort floors. Honor explicit model, effort, budget,
50
+ and data-processing constraints. Adapt when capabilities or evidence change.
51
+
52
+ ## 3. Make additional contexts earn their cost
53
+
54
+ Add a worker for a concrete benefit: needed expertise or tools, useful context or
55
+ permission isolation, independent scrutiny, or genuinely parallel work that
56
+ shortens the critical path after startup and integration costs. Compare this
57
+ with direct or batched tool use and continuing an existing worker.
58
+
59
+ Choose coherent, independently ownable units. Bundle related work that shares
60
+ context and acceptance criteria instead of splitting by file, phase, or role.
61
+ Independence makes parallel work possible; its expected contribution makes it
62
+ worthwhile. Each additional worker should provide a distinct useful result.
63
+
64
+ Parallel changes need clear interfaces and non-overlapping ownership or suitable
65
+ isolation, including shared test resources and integration assumptions. Use the
66
+ main context for complementary work rather than repeating a delegated task.
67
+ Apply the same benefit test to nested delegation and account for its total cost.
68
+
69
+ Several perspectives can be considered in one context. Separate them when
70
+ independent scrutiny, different evidence, or specialized capability adds value,
71
+ not merely to fill a persona panel. Simulated perspectives are not independent
72
+ review or observed user feedback.
73
+
74
+ ## 4. Transfer sufficient context, not entire histories
75
+
76
+ Reuse valid findings and existing worker contexts for related follow-up when
77
+ supported. Start fresh when independence, context overload, or a changed purpose
78
+ makes reuse counterproductive.
79
+
80
+ Give a worker the objective, essential constraints, relevant artifact and evidence
81
+ references, unresolved question, allowed changes, and completion criteria. Supply
82
+ enough context to decide correctly, including material failed approaches; favor
83
+ accessible source references over replaying the conversation. Keep critical
84
+ constraints explicit rather than relying on a lossy summary or hidden context.
85
+
86
+ Ask for the result needed for integration, its evidence, and remaining gaps.
87
+ Short results can return directly. Use durable files or artifacts for large,
88
+ reusable, or truncation-prone output, agreeing on the destination before it is
89
+ produced. Preserve relevant version information so evidence stays attributable.
90
+
91
+ ## 5. Review consequential uncertainty, not every stage
92
+
93
+ Implementers should execute relevant checks and correct their work. This is
94
+ verification, but not independent review. Add independent review when a fresh
95
+ perspective materially improves assurance or when explicitly required.
96
+
97
+ Give the reviewer the original objective, constraints, actual artifacts, and
98
+ execution evidence, with sufficient capability to assess the material risks.
99
+ Focus independent assessment on consequential assumptions and blind spots rather
100
+ than replaying the author's reasoning or assigning a reviewer to each artifact.
101
+ Neither model prestige nor agreement among agents substitutes for evidence.
102
+
103
+ Own acceptance and integration: inspect the artifacts and evidence material to
104
+ the outcome, resolve substantiated gaps, and check affected integration behavior.
105
+ Reuse applicable execution and review evidence across stages; revisit areas when
106
+ changes or new information invalidate it. Finish when required readiness is
107
+ supported, rather than commissioning another pass without a likely benefit.
108
+
109
+ ## 6. Preserve authority and continuity
110
+
111
+ Resolve model availability, effort controls, inheritance, and tool behavior from
112
+ the actual environment or applicable documentation when needed; reuse current
113
+ evidence. Distinguish requested settings from observed execution. Do not claim
114
+ a model, effort level, or independent review that was not established.
115
+
116
+ Apply the same scope, data protections, and approval boundaries to every worker,
117
+ whether native or external. Confirm that the actual provider, disclosed data,
118
+ and permitted actions are covered by existing authorization before routing.
119
+ Obtain approval when that boundary changes; an installed tool is not permission
120
+ to disclose repository content or apply changes to shared systems.
121
+
122
+ When a capability is unavailable, continue through an authorized alternative
123
+ that still meets required quality, reporting material substitutions or limits.
124
+ If a required gate cannot be satisfied, pause only dependent actions.
125
+
126
+ For interruptions or ownership changes, preserve the current state, artifact
127
+ locations, verified evidence, unresolved issues, and next action. Resume useful
128
+ contexts when supported. Collect results and release workers that no longer have
129
+ useful pending work, preserving required evidence and existing user work.
@@ -131,7 +131,7 @@ One shape means a worker never has to infer where the boundaries are stated, and
131
131
  delegation missing a field is visibly missing it.
132
132
 
133
133
  Where the `model-orchestration` skill is installed, its delegation-prompt spec is the routing-side
134
- authority (which model, which effort floor, who verifies). This template **carries** that spec
134
+ authority (which model and effort, whether another context earns its cost, who verifies). This template **carries** that spec
135
135
  rather than competing with it — the mapping is one-to-one:
136
136
 
137
137
  | Delegation element | Brief field |
@@ -148,8 +148,8 @@ Two of those rows are the ones people drop, so state them explicitly:
148
148
 
149
149
  - **Resource limit is where model and effort live.** "One agent, opus, effort floor xhigh, no
150
150
  further fan-out" belongs here as a hard cap, not as an aside in prose. A cap that is not in the
151
- brief is not a cap. Where `model-orchestration` is installed, take the floors from it; where it
152
- is not, still name the model and effort you intend, because a worker that inherits an unstated
151
+ brief is not a cap. Where `model-orchestration` is installed, take the model and effort choice
152
+ from its allocation principles; where it is not, still name the model and effort you intend, because a worker that inherits an unstated
153
153
  level runs at whatever the session happened to be set to.
154
154
  - **리뷰어 위임형 is the lane split, written down.** Choosing it tells the worker that a different
155
155
  lane writes the tests and issues the verdict, so it must not spend the run self-certifying —