@uzysjung/agent-harness 26.151.0 → 26.152.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.ko.md +2 -2
- package/README.md +2 -2
- package/dist/{chunk-3QBHZUVB.js → chunk-EQDC2AAU.js} +30 -23
- package/dist/chunk-EQDC2AAU.js.map +1 -0
- package/dist/index.js +133 -263
- package/dist/index.js.map +1 -1
- package/dist/trust-tier-drift.js +1 -1
- package/package.json +1 -1
- package/templates/rules/change-management.md +2 -3
- package/templates/rules/cli-development.md +2 -4
- package/templates/rules/doc-governance.md +1 -1
- package/templates/skills/audit-service-gaps/SKILL.md +0 -2
- package/templates/skills/model-orchestration/SKILL.md +6 -24
- package/templates/track-mcp-map.tsv +1 -1
- package/dist/chunk-3QBHZUVB.js.map +0 -1
- package/templates/agents/build-error-resolver.md +0 -111
- package/templates/agents/plan-checker.md +0 -116
- package/templates/agents/silent-failure-hunter.md +0 -50
- package/templates/skills/agent-introspection-debugging/SKILL.md +0 -153
- package/templates/skills/deep-research/SKILL.md +0 -179
- package/templates/skills/deep-research/agents/openai.yaml +0 -7
- package/templates/skills/eval-harness/SKILL.md +0 -306
- package/templates/skills/eval-harness/agents/openai.yaml +0 -7
- package/templates/skills/verification-loop/SKILL.md +0 -250
- package/templates/skills/verification-loop/agents/openai.yaml +0 -7
- package/templates/skills/verification-loop/references/tracks.md +0 -49
package/dist/trust-tier-drift.js
CHANGED
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@uzysjung/agent-harness",
|
|
3
|
-
"version": "26.
|
|
3
|
+
"version": "26.152.0",
|
|
4
4
|
"description": "Curate vetted AI-coding skills & plugins by your tech stack — install only what you need, across Claude Code, Codex, OpenCode & Antigravity",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"publishConfig": {
|
|
@@ -1,6 +1,5 @@
|
|
|
1
1
|
# Change Boundaries
|
|
2
2
|
|
|
3
|
-
- 완료 기준의 Pass/Fail · 다른 단계의 입출력 · 명시된 Non-Goals · 수정 금지 영역을 바꿔야 한다면 **임의로 바꾸지 않고 인간 결정을 받는다.**
|
|
4
|
-
-
|
|
5
|
-
- 스펙에 없던 결정 중 **아키텍처 · 외부 의존성 · 데이터 모델 · 보안 정책 · breaking API** 처럼 이후 작업에 계속 영향을 주는 것은 결정 기록으로 남긴다. 한 함수의 구현 디테일, 임시 워크어라운드, 명백한 버그 fix 는 대상이 아니다.
|
|
3
|
+
- 완료 기준의 Pass/Fail · 다른 단계의 입출력 · 명시된 Non-Goals · 수정 금지 영역을 바꿔야 한다면 **임의로 바꾸지 않고 인간 결정을 받는다.**
|
|
4
|
+
- 스펙에 없던 결정 중 **아키텍처 · 외부 의존성 · 데이터 모델 · 보안 정책 · breaking API** 처럼 이후 작업에 계속 영향을 주는 것은 결정 기록으로 남긴다. 한 함수의 구현 디테일, 임시 워크어라운드, 명백한 버그 fix 는 대상이 아니다. 이미 합의된 내용의 구체화도 기록을 남긴다.
|
|
6
5
|
- 결정을 바꿀 때는 **이전 기록의 상태와 그것을 참조하는 문서도 함께** 갱신한다. 한쪽만 고치면 어느 것이 현행인지 알 수 없다.
|
|
@@ -6,9 +6,7 @@ paths:
|
|
|
6
6
|
|
|
7
7
|
# Shell Safety
|
|
8
8
|
|
|
9
|
-
-
|
|
10
|
-
-
|
|
11
|
-
- 부정 결론("없다"·"안 된다")의 대조군 요구와 그것을 강제하는 도구는 `doc-governance` 에 있다.
|
|
12
|
-
- macOS(BSD)와 Linux(GNU)는 `sed -i` · `date` · `readlink -f` · `realpath -m` · `find -newermt` · `stat` 포맷이 호환되지 않는다. 양쪽에서 도는 형태를 쓰거나 `command -v` 로 분기한다.
|
|
9
|
+
- 부정 결론("없다"·"안 된다")은 `doc-governance` 가 지목한 `check-absence.sh` 로 낸다 — 대조군을 요구하고 stderr 를 보존한다.
|
|
10
|
+
- macOS(BSD)와 Linux(GNU)에서 다르게 도는 명령은 양쪽에서 도는 형태로 쓰거나 `command -v` 로 분기한다.
|
|
13
11
|
|
|
14
12
|
훅으로 쓸 스크립트의 **차단 계약은 실행기마다 다르다** — Claude Code · Codex 는 `exit 2` + stderr 에 사유(통과 = `exit 0`, 출력 없음) · OpenCode 플러그인은 훅 함수에서 `throw` · Antigravity 는 JSON 으로 `decision: "deny"`. 계약 밖의 형태는 "비차단 오류"로 흘러가 조용히 무시된다.
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
- **한 사실의 기준 문서는 하나다.** 다른 문서는 같은 내용을 독립적으로 복제하지 말고 가리키거나 필요한 만큼만 요약한다.
|
|
4
4
|
- 그 변경을 추적하는 문서가 있으면 **머지까지** 코드와 추적 상태를 함께 동기화한다. 한쪽만 바꾸면 다음 세션이 완료분을 중복 작업하거나 미완분을 완료로 오인해 건너뛴다.
|
|
5
|
-
-
|
|
5
|
+
- 추적 문서의 미완 표기는 코드로 확인해 정정한다 — 심볼이 있다는 것만으로 완료로 읽지 않는다.
|
|
6
6
|
|
|
7
7
|
- **대조군 없이 "없다"·"안 된다"를 결론으로 쓰지 마라.** 빈 결과도 실패한 명령도 "대상이 없다"와 "내 탐지기가 틀렸다"를 구분해 주지 않는다.
|
|
8
8
|
|
|
@@ -231,8 +231,6 @@ that skip benchmarking — is in [references/worked-example.md](references/worke
|
|
|
231
231
|
evaluation here.
|
|
232
232
|
- **`north-star`** — same target state, opposite direction: it *directs* the roadmap forward; this
|
|
233
233
|
*detects* gaps against it. It also owns the roadmap ordering this skill's findings feed.
|
|
234
|
-
- **`verification-loop`** — after a fix lands, it produces the evidence and the fixed verdict;
|
|
235
|
-
`VERIFY` mode consumes that rather than re-deriving it.
|
|
236
234
|
- **ADR conventions** — record each proposed fix as an architecture decision record in the
|
|
237
235
|
project's `docs/decisions/`, including the rejected benchmark alternative.
|
|
238
236
|
|
|
@@ -173,16 +173,9 @@ frontmatter model choice. If delegation models look wrong, check that env var fi
|
|
|
173
173
|
|
|
174
174
|
## Delegation prompt spec
|
|
175
175
|
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
|
|
179
|
-
1. **Objective** — one sentence, plus why it matters (models perform better knowing intent).
|
|
180
|
-
2. **Output format** — what comes back, in what structure. State that the final message IS the
|
|
181
|
-
deliverable (raw data, no user-facing preamble).
|
|
182
|
-
3. **Tool/source guidance** — where to look, what to trust, what to skip.
|
|
183
|
-
4. **Boundaries** — what NOT to touch, where the task ends, what is out of scope.
|
|
184
|
-
5. **Acceptance criteria** — how the worker (and you) know it's done. Strong AC lets the
|
|
185
|
-
worker loop independently instead of returning half-done.
|
|
176
|
+
[[objective-brief]] owns the delegation prompt format — write the brief there and hand it over.
|
|
177
|
+
Without that skill installed, the five elements are: objective (and why it matters), output
|
|
178
|
+
format, tool/source guidance, boundaries, acceptance criteria.
|
|
186
179
|
|
|
187
180
|
Scale worker count to the task, stated up front in your own plan: trivial lookup → no agent at
|
|
188
181
|
all (do it directly); bounded question → one agent; genuinely independent axes → one agent per
|
|
@@ -191,20 +184,9 @@ turn — delegate when the task's value justifies it, not by reflex.
|
|
|
191
184
|
|
|
192
185
|
## Worker lifecycle — 다 쓴 에이전트는 닫는다
|
|
193
186
|
|
|
194
|
-
|
|
195
|
-
|
|
196
|
-
|
|
197
|
-
after their job ended. The clutter compounds per delegation.
|
|
198
|
-
|
|
199
|
-
- **Consume the result → stop the agent** (TaskStop or the harness's stop mechanism) as one
|
|
200
|
-
motion. Close-after-use is the default, not a cleanup chore for later.
|
|
201
|
-
- Keep a worker alive ONLY when you will genuinely continue it via SendMessage (e.g. a reviewer
|
|
202
|
-
that must re-verify after fixes land) — and state that intent when you decide it, so every
|
|
203
|
-
still-open agent is a declared decision, not a leak.
|
|
204
|
-
- **Sweep at checkpoints**: at phase end and during [[compaction-handoff]], list running agents
|
|
205
|
-
and stop every finished one before moving on. Compaction wipes the orchestrator's own ledger
|
|
206
|
-
of open workers, so the sweep must use an enumeration probe (a TaskStop against a nonexistent
|
|
207
|
-
id lists running teammates), not memory.
|
|
187
|
+
Closing what you spawned is owned by the `git-policy` rule's Session Cleanup section — the
|
|
188
|
+
sweep there uses an enumeration probe rather than memory, because compaction wipes the
|
|
189
|
+
orchestrator's own ledger of open workers.
|
|
208
190
|
|
|
209
191
|
### Collect results as a file, not as a return message
|
|
210
192
|
|