@uzysjung/agent-harness 26.152.0 → 26.154.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/{chunk-EQDC2AAU.js → chunk-242TVGMZ.js} +582 -556
- package/dist/chunk-242TVGMZ.js.map +1 -0
- package/dist/index.js +623 -303
- package/dist/index.js.map +1 -1
- package/dist/trust-tier-drift.js +1 -1
- package/package.json +1 -1
- package/templates/CLAUDE.md +93 -145
- package/templates/codex/README.md +6 -7
- package/templates/codex/config.toml.template +1 -9
- package/templates/skills/audit-harness-fit/README.md +76 -43
- package/templates/skills/audit-harness-fit/SKILL.md +62 -47
- package/templates/skills/audit-harness-fit/evals/scenarios.yaml +179 -22
- package/templates/skills/audit-harness-fit/references/apply.md +74 -45
- package/templates/skills/audit-harness-fit/references/audit.md +142 -118
- package/templates/skills/audit-harness-fit/references/populate.md +53 -49
- package/templates/skills/audit-harness-fit/references/verification.md +131 -100
- package/templates/skills/clear-korean-communication/SKILL.md +76 -209
- package/templates/skills/external-model-consult/scripts/codex-ask.sh +9 -4
- package/templates/skills/external-model-consult/scripts/gemini-ask.sh +9 -4
- package/templates/skills/model-orchestration/SKILL.md +123 -267
- package/templates/skills/objective-brief/SKILL.md +3 -3
- package/dist/chunk-EQDC2AAU.js.map +0 -1
- package/templates/codex/hooks/README.md +0 -37
- package/templates/codex/hooks/session-start.sh +0 -7
- package/templates/codex/hooks/uncommitted-check.sh +0 -7
- package/templates/skills/clear-korean-communication/references/pre-send-checklist.md +0 -37
- package/templates/skills/clear-korean-communication/references/why-it-works.md +0 -24
- package/templates/skills/clear-korean-communication/references/worked-examples.md +0 -89
|
@@ -146,17 +146,22 @@ env -i PATH="$PATH" HOME="$HOME" TERM="${TERM:-dumb}" \
|
|
|
146
146
|
"$CODEX" "${ARGS[@]}" 1>&2 &
|
|
147
147
|
CPID=$!
|
|
148
148
|
TIMED_OUT=0
|
|
149
|
-
|
|
149
|
+
# Poll at 200 ms, not 1 s (#411): with a 1 s tick every call paid up to a full second of idle
|
|
150
|
+
# wait after the CLI had already exited, which is what pushed the wrapper's test cases toward
|
|
151
|
+
# vitest's 5 s budget on a busy machine. TICKS counts fifths of a second; both BSD and GNU
|
|
152
|
+
# `sleep` accept the fractional argument.
|
|
153
|
+
TICKS=0
|
|
154
|
+
TICK_LIMIT=$((TIMEOUT_S * 5))
|
|
150
155
|
while kill -0 "$CPID" 2>/dev/null; do
|
|
151
|
-
if [ "$
|
|
156
|
+
if [ "$TICKS" -ge "$TICK_LIMIT" ]; then
|
|
152
157
|
TIMED_OUT=1
|
|
153
158
|
kill "$CPID" 2>/dev/null || true
|
|
154
159
|
sleep 2
|
|
155
160
|
kill -9 "$CPID" 2>/dev/null || true
|
|
156
161
|
break
|
|
157
162
|
fi
|
|
158
|
-
sleep
|
|
159
|
-
|
|
163
|
+
sleep 0.2
|
|
164
|
+
TICKS=$((TICKS + 1))
|
|
160
165
|
done
|
|
161
166
|
set +e
|
|
162
167
|
wait "$CPID"
|
|
@@ -183,17 +183,22 @@ env -i PATH="$PATH" HOME="$HOME" TERM="${TERM:-dumb}" \
|
|
|
183
183
|
"$AGY" --model="$MODEL" -p "$PROMPT" >"$MSGFILE" &
|
|
184
184
|
APID=$!
|
|
185
185
|
TIMED_OUT=0
|
|
186
|
-
|
|
186
|
+
# Poll at 200 ms, not 1 s (#411): with a 1 s tick every call paid up to a full second of idle
|
|
187
|
+
# wait after the CLI had already exited, which is what pushed the wrapper's test cases toward
|
|
188
|
+
# vitest's 5 s budget on a busy machine. TICKS counts fifths of a second; both BSD and GNU
|
|
189
|
+
# `sleep` accept the fractional argument.
|
|
190
|
+
TICKS=0
|
|
191
|
+
TICK_LIMIT=$((TIMEOUT_S * 5))
|
|
187
192
|
while kill -0 "$APID" 2>/dev/null; do
|
|
188
|
-
if [ "$
|
|
193
|
+
if [ "$TICKS" -ge "$TICK_LIMIT" ]; then
|
|
189
194
|
TIMED_OUT=1
|
|
190
195
|
kill "$APID" 2>/dev/null || true
|
|
191
196
|
sleep 2
|
|
192
197
|
kill -9 "$APID" 2>/dev/null || true
|
|
193
198
|
break
|
|
194
199
|
fi
|
|
195
|
-
sleep
|
|
196
|
-
|
|
200
|
+
sleep 0.2
|
|
201
|
+
TICKS=$((TICKS + 1))
|
|
197
202
|
done
|
|
198
203
|
set +e
|
|
199
204
|
wait "$APID"
|
|
@@ -1,273 +1,129 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: model-orchestration
|
|
3
3
|
description: >-
|
|
4
|
-
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
repetitive/simple implementation goes to Sonnet at high or above — Sonnet is never used for
|
|
10
|
-
tests or verification; plan/spec drafts may be produced by Opus from a Fable direction brief,
|
|
11
|
-
but the review/decision is always Fable's. Never delegate below the effort floors. Use whenever
|
|
12
|
-
you are about to spawn an Agent/Task/Workflow worker, pick a model for a subtask, set a
|
|
13
|
-
thinking/effort level, or assign verification. Trigger on "위임해", "오케스트레이션", "모델 역할분담", "어떤 모델로",
|
|
14
|
-
"effort 얼마로", "서브에이전트", "delegate this", "which model should". Fire even when the
|
|
15
|
-
user doesn't name the policy — any delegation is in scope.
|
|
4
|
+
Select models, reasoning effort, and delegation boundaries for a task. Use when
|
|
5
|
+
choosing direct execution versus subagents, allocating parallel work or
|
|
6
|
+
independent review, or reducing multi-agent token and context overhead. Favor
|
|
7
|
+
context reuse and add model calls or workers when their expected contribution
|
|
8
|
+
justifies the cost.
|
|
16
9
|
---
|
|
17
10
|
|
|
18
|
-
# Model Orchestration
|
|
11
|
+
# Model Orchestration
|
|
19
12
|
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
the
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
The delegation prompt spec and the file-handoff contract below apply here unchanged: an outside
|
|
138
|
-
worker is still a worker.
|
|
139
|
-
|
|
140
|
-
## Effort floors — and the inheritance gotcha
|
|
141
|
-
|
|
142
|
-
Effort tiers: `low` < `medium` < `high` < `xhigh` < `max`. Fable 5, Sonnet 5, and Opus 4.8/4.7
|
|
143
|
-
support all five; on models without `xhigh` the request silently falls back to the nearest
|
|
144
|
-
lower level — check the model before assuming the floor holds.
|
|
145
|
-
|
|
146
|
-
**The gotcha that breaks this policy silently:** the Agent/Task tool accepts a per-invocation
|
|
147
|
-
`model`, but **not a per-invocation `effort`** — a spawned agent inherits the session effort
|
|
148
|
-
unless its definition says otherwise. Session default is `high`, so a plain
|
|
149
|
-
`Agent(model: "opus")` runs at `high`, **below the xhigh floor**, with no warning. Enforce the
|
|
150
|
-
floor through one of these three paths:
|
|
151
|
-
|
|
152
|
-
1. **Pre-defined agent roles** — `.claude/agents/<role>.md` frontmatter pins both knobs.
|
|
153
|
-
This is the durable path for recurring roles:
|
|
154
|
-
|
|
155
|
-
```yaml
|
|
156
|
-
---
|
|
157
|
-
name: verifier
|
|
158
|
-
description: Fresh-context V&V per model-orchestration policy
|
|
159
|
-
model: opus
|
|
160
|
-
effort: xhigh
|
|
161
|
-
---
|
|
162
|
-
```
|
|
163
|
-
|
|
164
|
-
2. **Workflow scripts** — `agent(prompt, {model: "opus", effort: "xhigh"})` supports both
|
|
165
|
-
per call. Use for scripted fan-outs.
|
|
166
|
-
|
|
167
|
-
3. **Session inheritance** — if the session already runs at `xhigh` (e.g. `/effort xhigh` or
|
|
168
|
-
ultracode), a bare `Agent(model: "opus")` inherits a compliant level. Verify with
|
|
169
|
-
`/effort`; don't assume.
|
|
170
|
-
|
|
171
|
-
One environment caveat: `CLAUDE_CODE_SUBAGENT_MODEL` outranks every per-invocation and
|
|
172
|
-
frontmatter model choice. If delegation models look wrong, check that env var first.
|
|
173
|
-
|
|
174
|
-
## Delegation prompt spec
|
|
175
|
-
|
|
176
|
-
[[objective-brief]] owns the delegation prompt format — write the brief there and hand it over.
|
|
177
|
-
Without that skill installed, the five elements are: objective (and why it matters), output
|
|
178
|
-
format, tool/source guidance, boundaries, acceptance criteria.
|
|
179
|
-
|
|
180
|
-
Scale worker count to the task, stated up front in your own plan: trivial lookup → no agent at
|
|
181
|
-
all (do it directly); bounded question → one agent; genuinely independent axes → one agent per
|
|
182
|
-
axis. Over-spawning is a documented failure mode, and multi-agent runs cost ~15× a plain chat
|
|
183
|
-
turn — delegate when the task's value justifies it, not by reflex.
|
|
184
|
-
|
|
185
|
-
## Worker lifecycle — 다 쓴 에이전트는 닫는다
|
|
186
|
-
|
|
187
|
-
Closing what you spawned is owned by the `git-policy` rule's Session Cleanup section — the
|
|
188
|
-
sweep there uses an enumeration probe rather than memory, because compaction wipes the
|
|
189
|
-
orchestrator's own ledger of open workers.
|
|
190
|
-
|
|
191
|
-
### Collect results as a file, not as a return message
|
|
192
|
-
|
|
193
|
-
A worker's final message is not durable transport — long results get truncated silently, and
|
|
194
|
-
what was cut is gone. This burned three separate times — most
|
|
195
|
-
recently a panel run that lost three reviewers' output — before it was codified as a gate.
|
|
196
|
-
The contract: the worker writes its deliverable to a scratchpad/output **file whose path the
|
|
197
|
-
orchestrator fixes in the spawn prompt** — collection is a spawn-time decision, because a
|
|
198
|
-
truncated reply cannot be retrofit into a file after the fact. The final message then carries
|
|
199
|
-
only the file path plus a short summary. Treat any delegation whose prompt names no output
|
|
200
|
-
file as not yet dispatched.
|
|
201
|
-
|
|
202
|
-
## V&V separation
|
|
203
|
-
|
|
204
|
-
The implementer never verifies its own work — the *instance* that wrote something never judges
|
|
205
|
-
it. How that plays out per lane:
|
|
206
|
-
|
|
207
|
-
- **Sonnet implemented** (반복·단순 구현) → **Opus cross-verifies** — a different model has
|
|
208
|
-
different blind spots, which is the cheapest anchoring break available.
|
|
209
|
-
- **Opus implemented** (핵심 구현·테스트) → a **fresh Opus instance** verifies: a NEW agent with no
|
|
210
|
-
shared history. A verifier that watched the implementation happen inherits the implementer's
|
|
211
|
-
mental model and anchors on it — same-session self-review reliably misses the same edge
|
|
212
|
-
cases the implementation missed. Instance separation is what makes "Opus builds AND Opus
|
|
213
|
-
verifies" coherent.
|
|
214
|
-
- **Orchestrator layer on top**: the orchestrator hunts 성능·보안 문제점 in shipped features —
|
|
215
|
-
a second, higher-altitude pass that judges what a diff-level verifier doesn't (product fit,
|
|
216
|
-
systemic risk). Documents get the same split: Opus authors, the orchestrator reviews via
|
|
217
|
-
[[multi-persona-review]].
|
|
218
|
-
|
|
219
|
-
This pairs with, not replaces, deterministic gates (tests, typecheck, CI) — the verifier
|
|
220
|
-
judges what automation can't: spec fit, missed edge cases, design drift.
|
|
221
|
-
|
|
222
|
-
## Orchestrator handoff (quota exhaustion)
|
|
223
|
-
|
|
224
|
-
When the top-tier orchestrator's quota runs out mid-project, **Opus @ `max` takes over
|
|
225
|
-
orchestration**. There is no reliable automatic path — documented fallback chains explicitly
|
|
226
|
-
exclude rate-limit errors, and plan-level auto-switching is undocumented behavior you must not
|
|
227
|
-
build on. Hand off manually:
|
|
228
|
-
|
|
229
|
-
1. Run the [[compaction-handoff]] protocol: persist durable facts to memory, take an atomic
|
|
230
|
-
git snapshot (clean tree + open-PR check), emit the fixed-field resume anchor
|
|
231
|
-
(current state / verified / what's left / next action).
|
|
232
|
-
2. The successor session starts on `opus` at `max` (`/model opus` + `/effort max`), reads the
|
|
233
|
-
anchor, and continues as orchestrator under this same policy — setting direction, reviewing,
|
|
234
|
-
and hunting problems itself (it can author documents directly too, since it already runs at
|
|
235
|
-
the Opus doc-author tier — no separate delegation needed while it stands in).
|
|
236
|
-
3. When the top-tier model becomes available again, hand back the same way.
|
|
237
|
-
|
|
238
|
-
## Anti-patterns
|
|
239
|
-
|
|
240
|
-
| Anti-pattern | Why it's a violation |
|
|
241
|
-
|---|---|
|
|
242
|
-
| `Agent(model: "opus")` with session at default effort | Inherits `high` < xhigh floor — use a pinned agent role, Workflow opts, or raise session effort |
|
|
243
|
-
| Delegating 방향성 수립 or a final judgment call to a worker | Direction is the orchestrator's own — a delegation prompt can't carry the shaping context |
|
|
244
|
-
| Accepting an Opus-authored spec/plan without orchestrator review | Author≠reviewer applies to documents too — run [[multi-persona-review]] before accepting |
|
|
245
|
-
| Orchestrator hand-writing full spec/plan drafts itself | Authoring is Opus's lane — brief the direction, delegate the draft, review the result |
|
|
246
|
-
| Sonnet on tests, verification, or **core** implementation | Sonnet 의 레인은 반복·단순 구현(확립 패턴 적용)과 읽기 전용 보조까지다 — 테스트·V&V·새 설계 판단은 Opus/Fable 몫 |
|
|
247
|
-
| Sonnet 반복 구현물이 Opus 교차검증 없이 머지를 게이트 | 교차검증이 Sonnet 레인의 존립 조건이다 — 검증 없는 반복 구현은 레인 위반과 같다 |
|
|
248
|
-
| Effort below floor "to save tokens" | The floors ARE the policy; surface cost concerns to the user instead |
|
|
249
|
-
| Implementer verifying its own diff | Anchoring — verification needs fresh context, prefer a different model |
|
|
250
|
-
| Two agents writing the same files in parallel | Conflicting implicit decisions; keep writes sequential or worktree-isolated |
|
|
251
|
-
| Spawning an agent for what one direct tool call answers | 15× token multiplier for zero value — do trivial work directly |
|
|
252
|
-
| Relying on plan-level auto-fallback for continuity | Undocumented behavior; use the manual handoff protocol |
|
|
253
|
-
| Finished worker left running after its result is consumed | Subagent panes/windows accumulate (iTerm2 등) and idle pings pollute the session — TaskStop as one motion with consuming the result |
|
|
254
|
-
| 외부 실행기에 **테스트 작성·검증**·핵심 구현을 넘김 | 이 레인은 판단 잔여 0 인 일만 받는다 — 무엇을 단언할지 정하는 일을 밖으로 내보내면 외부 산출물을 검사할 기준 자체가 밖에 있게 된다. 형태가 이미 고정된 표에 케이스 한 줄을 복제하는 일은 무엇을 단언할지 정하지 않으므로 여기 해당하지 않는다 |
|
|
255
|
-
| 도구가 없어서 조용히 다른 제공자로 갈아타 실행 | 사용자는 어느 도구가 답했는지 알고 있다 — 대체는 보고 대상이지 판단 대상이 아니다 |
|
|
256
|
-
| 외부 실행기를 "품질이 더 낫다"는 이유로 고름 | 이 레인이 사는 것은 용량이다. 품질을 근거로 들면 이 정책의 quality-over-cost 전제를 뒤집는 것이다 |
|
|
257
|
-
|
|
258
|
-
## Quick reference
|
|
259
|
-
|
|
260
|
-
```
|
|
261
|
-
방향성 수립 / 설계·기획 확정 / 작업 분배 / 스펙 리뷰(multi-persona-review) / 기능 개선 / 성능·보안 문제발굴
|
|
262
|
-
→ 오케스트레이터(Fable) 직접
|
|
263
|
-
스펙 문서 초안(Fable 브리프 기반) / 핵심 구현 / 테스트 작성·실행(E2E 포함) / 코드 리뷰·검증(V&V)
|
|
264
|
-
→ opus @ xhigh (또는 max) — pinned role 또는 Workflow opts
|
|
265
|
-
반복·단순 구현(확립 패턴 적용) / 리서치 스윕
|
|
266
|
-
→ sonnet @ high 이상 — 테스트·검증 투입 금지, 산출물은 Opus 교차검증
|
|
267
|
-
결정적 변환 (rename·포맷) → 모델 위임 금지 — sed/grep/스크립트 직접
|
|
268
|
-
판단 잔여 0 + 합격을 명령 하나로 판정
|
|
269
|
-
→ sonnet 기본 / 다섯 술어 충족 시 외부 실행기 — 최초 1회 사용자 확인, 산출물은 in-harness 교차검증
|
|
270
|
-
도구 부재·인증 만료 → 레인을 내리고 무엇을 못 썼는지 보고 — 대신 설치/로그인 금지, 조용한 제공자 교체 금지
|
|
271
|
-
위임 완료 → 결과 수거와 동시에 TaskStop (SendMessage 재사용 예정 시만 유지 선언)
|
|
272
|
-
Fable 소진 → compaction-handoff → opus @ max 가 오케스트레이터 대행
|
|
273
|
-
```
|
|
13
|
+
Apply intelligence where it most improves a verified, usable outcome. Optimize
|
|
14
|
+
whole-task cost and delay, including context duplication, coordination,
|
|
15
|
+
verification, and likely rework, while preserving required quality and authority.
|
|
16
|
+
These are allocation principles, not a mandatory multi-agent workflow. Explicit
|
|
17
|
+
project requirements remain binding.
|
|
18
|
+
|
|
19
|
+
## 1. Keep ownership with useful context
|
|
20
|
+
|
|
21
|
+
The agent holding the user's intent and relevant project context owns the
|
|
22
|
+
integrated result. It may plan, draft, implement, run checks, and refine directly.
|
|
23
|
+
Roles describe responsibilities; they do not require separate workers.
|
|
24
|
+
|
|
25
|
+
Keep tightly coupled work together when a handoff would mostly make another
|
|
26
|
+
agent reconstruct the same understanding. Completing the task in one capable
|
|
27
|
+
context is a valid orchestration choice, not a failure to delegate.
|
|
28
|
+
|
|
29
|
+
Treat model capability, reasoning effort, and context count as separate choices.
|
|
30
|
+
Where supported, changing model or effort within an existing context may be more
|
|
31
|
+
useful than spawning another agent. Reuse applicable routing decisions rather
|
|
32
|
+
than reopening them for every tool call or phase.
|
|
33
|
+
|
|
34
|
+
## 2. Put capability at the bottleneck
|
|
35
|
+
|
|
36
|
+
Use stronger reasoning for unresolved design choices, difficult diagnosis,
|
|
37
|
+
cross-boundary behavior, and consequential mistakes. Choose sufficient capability
|
|
38
|
+
early when uncertainty or recovery cost warrants it; there is no need to exhaust
|
|
39
|
+
weaker models first. A stronger model doing the work directly may cost less
|
|
40
|
+
overall than a weaker implementation followed by repeated review and repair.
|
|
41
|
+
|
|
42
|
+
For well-specified work, choose between direct execution, deterministic tools,
|
|
43
|
+
and a lighter capable model by total completion cost. Running a known test command
|
|
44
|
+
and deciding whether its coverage is adequate are different demands on judgment.
|
|
45
|
+
Use tools for execution and measurement; allocate reasoning to decisions they
|
|
46
|
+
cannot settle.
|
|
47
|
+
|
|
48
|
+
Choose from actual capabilities and task evidence rather than fixed vendor-role
|
|
49
|
+
assignments or universal effort floors. Honor explicit model, effort, budget,
|
|
50
|
+
and data-processing constraints. Adapt when capabilities or evidence change.
|
|
51
|
+
|
|
52
|
+
## 3. Make additional contexts earn their cost
|
|
53
|
+
|
|
54
|
+
Add a worker for a concrete benefit: needed expertise or tools, useful context or
|
|
55
|
+
permission isolation, independent scrutiny, or genuinely parallel work that
|
|
56
|
+
shortens the critical path after startup and integration costs. Compare this
|
|
57
|
+
with direct or batched tool use and continuing an existing worker.
|
|
58
|
+
|
|
59
|
+
Choose coherent, independently ownable units. Bundle related work that shares
|
|
60
|
+
context and acceptance criteria instead of splitting by file, phase, or role.
|
|
61
|
+
Independence makes parallel work possible; its expected contribution makes it
|
|
62
|
+
worthwhile. Each additional worker should provide a distinct useful result.
|
|
63
|
+
|
|
64
|
+
Parallel changes need clear interfaces and non-overlapping ownership or suitable
|
|
65
|
+
isolation, including shared test resources and integration assumptions. Use the
|
|
66
|
+
main context for complementary work rather than repeating a delegated task.
|
|
67
|
+
Apply the same benefit test to nested delegation and account for its total cost.
|
|
68
|
+
|
|
69
|
+
Several perspectives can be considered in one context. Separate them when
|
|
70
|
+
independent scrutiny, different evidence, or specialized capability adds value,
|
|
71
|
+
not merely to fill a persona panel. Simulated perspectives are not independent
|
|
72
|
+
review or observed user feedback.
|
|
73
|
+
|
|
74
|
+
## 4. Transfer sufficient context, not entire histories
|
|
75
|
+
|
|
76
|
+
Reuse valid findings and existing worker contexts for related follow-up when
|
|
77
|
+
supported. Start fresh when independence, context overload, or a changed purpose
|
|
78
|
+
makes reuse counterproductive.
|
|
79
|
+
|
|
80
|
+
Give a worker the objective, essential constraints, relevant artifact and evidence
|
|
81
|
+
references, unresolved question, allowed changes, and completion criteria. Supply
|
|
82
|
+
enough context to decide correctly, including material failed approaches; favor
|
|
83
|
+
accessible source references over replaying the conversation. Keep critical
|
|
84
|
+
constraints explicit rather than relying on a lossy summary or hidden context.
|
|
85
|
+
|
|
86
|
+
Ask for the result needed for integration, its evidence, and remaining gaps.
|
|
87
|
+
Short results can return directly. Use durable files or artifacts for large,
|
|
88
|
+
reusable, or truncation-prone output, agreeing on the destination before it is
|
|
89
|
+
produced. Preserve relevant version information so evidence stays attributable.
|
|
90
|
+
|
|
91
|
+
## 5. Review consequential uncertainty, not every stage
|
|
92
|
+
|
|
93
|
+
Implementers should execute relevant checks and correct their work. This is
|
|
94
|
+
verification, but not independent review. Add independent review when a fresh
|
|
95
|
+
perspective materially improves assurance or when explicitly required.
|
|
96
|
+
|
|
97
|
+
Give the reviewer the original objective, constraints, actual artifacts, and
|
|
98
|
+
execution evidence, with sufficient capability to assess the material risks.
|
|
99
|
+
Focus independent assessment on consequential assumptions and blind spots rather
|
|
100
|
+
than replaying the author's reasoning or assigning a reviewer to each artifact.
|
|
101
|
+
Neither model prestige nor agreement among agents substitutes for evidence.
|
|
102
|
+
|
|
103
|
+
Own acceptance and integration: inspect the artifacts and evidence material to
|
|
104
|
+
the outcome, resolve substantiated gaps, and check affected integration behavior.
|
|
105
|
+
Reuse applicable execution and review evidence across stages; revisit areas when
|
|
106
|
+
changes or new information invalidate it. Finish when required readiness is
|
|
107
|
+
supported, rather than commissioning another pass without a likely benefit.
|
|
108
|
+
|
|
109
|
+
## 6. Preserve authority and continuity
|
|
110
|
+
|
|
111
|
+
Resolve model availability, effort controls, inheritance, and tool behavior from
|
|
112
|
+
the actual environment or applicable documentation when needed; reuse current
|
|
113
|
+
evidence. Distinguish requested settings from observed execution. Do not claim
|
|
114
|
+
a model, effort level, or independent review that was not established.
|
|
115
|
+
|
|
116
|
+
Apply the same scope, data protections, and approval boundaries to every worker,
|
|
117
|
+
whether native or external. Confirm that the actual provider, disclosed data,
|
|
118
|
+
and permitted actions are covered by existing authorization before routing.
|
|
119
|
+
Obtain approval when that boundary changes; an installed tool is not permission
|
|
120
|
+
to disclose repository content or apply changes to shared systems.
|
|
121
|
+
|
|
122
|
+
When a capability is unavailable, continue through an authorized alternative
|
|
123
|
+
that still meets required quality, reporting material substitutions or limits.
|
|
124
|
+
If a required gate cannot be satisfied, pause only dependent actions.
|
|
125
|
+
|
|
126
|
+
For interruptions or ownership changes, preserve the current state, artifact
|
|
127
|
+
locations, verified evidence, unresolved issues, and next action. Resume useful
|
|
128
|
+
contexts when supported. Collect results and release workers that no longer have
|
|
129
|
+
useful pending work, preserving required evidence and existing user work.
|
|
@@ -131,7 +131,7 @@ One shape means a worker never has to infer where the boundaries are stated, and
|
|
|
131
131
|
delegation missing a field is visibly missing it.
|
|
132
132
|
|
|
133
133
|
Where the `model-orchestration` skill is installed, its delegation-prompt spec is the routing-side
|
|
134
|
-
authority (which model
|
|
134
|
+
authority (which model and effort, whether another context earns its cost, who verifies). This template **carries** that spec
|
|
135
135
|
rather than competing with it — the mapping is one-to-one:
|
|
136
136
|
|
|
137
137
|
| Delegation element | Brief field |
|
|
@@ -148,8 +148,8 @@ Two of those rows are the ones people drop, so state them explicitly:
|
|
|
148
148
|
|
|
149
149
|
- **Resource limit is where model and effort live.** "One agent, opus, effort floor xhigh, no
|
|
150
150
|
further fan-out" belongs here as a hard cap, not as an aside in prose. A cap that is not in the
|
|
151
|
-
brief is not a cap. Where `model-orchestration` is installed, take the
|
|
152
|
-
is not, still name the model and effort you intend, because a worker that inherits an unstated
|
|
151
|
+
brief is not a cap. Where `model-orchestration` is installed, take the model and effort choice
|
|
152
|
+
from its allocation principles; where it is not, still name the model and effort you intend, because a worker that inherits an unstated
|
|
153
153
|
level runs at whatever the session happened to be set to.
|
|
154
154
|
- **리뷰어 위임형 is the lane split, written down.** Choosing it tells the worker that a different
|
|
155
155
|
lane writes the tests and issues the verdict, so it must not spend the run self-certifying —
|