@uzysjung/agent-harness 26.150.0 → 26.151.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.ko.md +1 -1
- package/README.md +1 -1
- package/dist/{chunk-NKBUDHPC.js → chunk-3QBHZUVB.js} +99 -50
- package/dist/chunk-3QBHZUVB.js.map +1 -0
- package/dist/index.js +389 -274
- package/dist/index.js.map +1 -1
- package/dist/trust-tier-drift.js +1 -1
- package/package.json +1 -1
- package/templates/CLAUDE.md +145 -164
- package/templates/agents/build-error-resolver.md +1 -1
- package/templates/agents/plan-checker.md +1 -1
- package/templates/agents/reviewer.md +4 -5
- package/templates/antigravity/AGENTS.md.template +3 -23
- package/templates/codex/AGENTS.md.template +5 -56
- package/templates/hooks/protect-files.sh +4 -0
- package/templates/opencode/AGENTS.md.template +4 -52
- package/templates/opencode/opencode.json.template +0 -8
- package/templates/rules/change-management.md +0 -1
- package/templates/rules/cli-development.md +1 -1
- package/templates/rules/doc-governance.md +2 -0
- package/templates/rules/git-policy.md +1 -1
- package/templates/rules/ship-checklist.md +3 -3
- package/templates/rules/test-policy.md +3 -8
- package/templates/settings.json +1 -16
- package/templates/skills/agent-introspection-debugging/SKILL.md +1 -1
- package/templates/skills/audit-harness-fit/README.md +113 -0
- package/templates/skills/audit-harness-fit/SKILL.md +64 -433
- package/templates/skills/audit-harness-fit/evals/scenarios.yaml +222 -0
- package/templates/skills/audit-harness-fit/references/apply.md +66 -0
- package/templates/skills/audit-harness-fit/references/audit.md +160 -0
- package/templates/skills/audit-harness-fit/references/populate.md +74 -0
- package/templates/skills/audit-harness-fit/references/verification.md +123 -0
- package/templates/skills/compaction-handoff/SKILL.md +2 -2
- package/templates/skills/model-orchestration/SKILL.md +7 -0
- package/templates/skills/natural-korean/SKILL.md +45 -0
- package/templates/skills/north-star/references/roadmap-method.md +2 -2
- package/templates/skills/{task-brief → objective-brief}/SKILL.md +17 -16
- package/templates/skills/recurrence-prevention/SKILL.md +16 -16
- package/dist/chunk-NKBUDHPC.js.map +0 -1
- package/templates/agents/code-reviewer.md +0 -237
- package/templates/agents/security-reviewer.md +0 -108
- package/templates/hooks/task-brief-nudge.sh +0 -57
- package/templates/skills/audit-harness-fit/references/official-criteria.md +0 -367
- package/templates/skills/continuous-learning-v2/SKILL.md +0 -361
- package/templates/skills/continuous-learning-v2/agents/observer-loop.sh +0 -362
- package/templates/skills/continuous-learning-v2/agents/observer.md +0 -189
- package/templates/skills/continuous-learning-v2/agents/session-guardian.sh +0 -150
- package/templates/skills/continuous-learning-v2/agents/start-observer.sh +0 -252
- package/templates/skills/continuous-learning-v2/config.json +0 -8
- package/templates/skills/continuous-learning-v2/hooks/observe.sh +0 -585
- package/templates/skills/continuous-learning-v2/scripts/detect-project.sh +0 -322
- package/templates/skills/continuous-learning-v2/scripts/instinct-cli.py +0 -1956
- package/templates/skills/continuous-learning-v2/scripts/lib/homunculus-dir.sh +0 -31
- package/templates/skills/continuous-learning-v2/scripts/migrate-homunculus.sh +0 -68
- package/templates/skills/continuous-learning-v2/scripts/test_parse_instinct.py +0 -1420
- package/templates/skills/humanize-korean/SKILL.md +0 -228
- package/templates/skills/spec-scaling/SKILL.md +0 -89
- package/templates/skills/strategic-compact/SKILL.md +0 -145
- package/templates/skills/strategic-compact/suggest-compact.sh +0 -54
|
@@ -22,7 +22,8 @@ This protocol is **snapshot-based, not append-based**:
|
|
|
22
22
|
- Git and PR state: authoritative implementation snapshot.
|
|
23
23
|
- `/compact` line: short pointer to the anchor, not another summary.
|
|
24
24
|
|
|
25
|
-
>
|
|
25
|
+
> Claude Code compacts on its own when the context fills up. This skill defines **how** to
|
|
26
|
+
> checkpoint before that happens — or before you compact by hand.
|
|
26
27
|
|
|
27
28
|
## Goals
|
|
28
29
|
|
|
@@ -325,5 +326,4 @@ a minimal manual resume anchor in the response.
|
|
|
325
326
|
|
|
326
327
|
## Related skills
|
|
327
328
|
|
|
328
|
-
- **strategic-compact** — decides when to compact.
|
|
329
329
|
- **git-policy Session Cleanup** — defines repository and PR cleanup expectations.
|
|
@@ -17,6 +17,13 @@ description: >-
|
|
|
17
17
|
|
|
18
18
|
# Model Orchestration Policy
|
|
19
19
|
|
|
20
|
+
**What is policy and what is vocabulary.** The policy is the role split — orchestrator ≠ builder ≠
|
|
21
|
+
verifier — and the rule that verification never goes to a lower tier than the work it verifies.
|
|
22
|
+
`Fable` · `Opus` · `Sonnet` and the effort levels (`xhigh` · `high` · `max`) are Claude Code
|
|
23
|
+
vocabulary. On Codex, OpenCode, or Antigravity apply the same split with that vendor's top tier for
|
|
24
|
+
orchestration and verification and its mid tier for repetitive implementation; an effort knob that
|
|
25
|
+
does not exist there is not a violation — putting verification on a lower tier is.
|
|
26
|
+
|
|
20
27
|
A fixed role split between model tiers, set by the user (2026-07-04, revised 2026-07-07 and
|
|
21
28
|
2026-08-02 — *"설계·분배·기획·리뷰는 Fable5, 핵심 구현·테스트·검증은 opus5, 반복·단순 구현은
|
|
22
29
|
sonnet"*). The premise is
|
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: natural-korean
|
|
3
|
+
description: 한국어 답변·글쓰기·번역·퇴고에 사용한다. 뜻과 말투를 살리면서 번역투·불필요한 영어·상투구·기계적 반복을 줄인다. "자연스럽게 다듬어줘", "AI 티 없애줘", "번역투 고쳐줘", "문체 진단" 요청에도 적용한다.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# 자연스러운 한국어
|
|
7
|
+
|
|
8
|
+
쉽고 익숙한 한국어로 뜻을 정확하게 전달한다. 사람처럼 보이게 꾸미기보다, 실제로 어색한 표현을 고친다.
|
|
9
|
+
|
|
10
|
+
## 작업 방식
|
|
11
|
+
|
|
12
|
+
- **새 글·답변:** 독자와 목적에 맞게 쓴다. 설명과 답변은 핵심부터 말하고, 별도 요청이 없는 대화는 자연스러운 존댓말을 쓴다.
|
|
13
|
+
- **번역:** 단어보다 뜻을 옮긴다. 정보와 뉘앙스는 유지하되, 어순·문장 구성·관용구는 자연스러운 한국어로 바꾼다.
|
|
14
|
+
- **퇴고:** 글 전체의 문맥을 파악한 뒤, 실제 문제를 진단하고 필요한 부분만 고친다. 별도 요청이 없으면 원래 말투·구조를 유지한다. 재작성이나 말투 변경을 요청하면 그 범위에서 바꾼다.
|
|
15
|
+
|
|
16
|
+
## 표현 원칙
|
|
17
|
+
|
|
18
|
+
1. **독자에게 익숙한 말을 쓴다.** 전문용어의 뜻은 보존하되, 설명문에서는 ‘스킬·규칙·프롬프트’처럼 익숙한 표기를 우선한다. 제품명·식별자·통용 약어는 유지한다. 낯선 용어는 필요할 때 짧게 설명하고, 원어 병기는 구분에 필요한 경우만 한다. 뜻이 불분명한 용어를 짐작으로 풀지 않는다.
|
|
19
|
+
2. **누가 무엇을 하는지 분명히 쓴다.** 명사를 길게 잇기보다 동사로 직접 표현한다. 불필요한 대명사·명사화·간접 표현을 덜되, 주체와 논리관계는 남긴다. 긴 문장은 필요할 때 나누고, 짧게 쓰려고 조사·어미·필요한 설명을 빼지 않는다.
|
|
20
|
+
3. **반복과 군더더기를 줄인다.** 내용에 보탬이 없는 칭찬·예고·과장·결론 반복은 덜어낸다. 접속사·완곡 표현·같은 문형도 문맥상 불필요하거나 반복이 거슬릴 때만 고친다. 반복을 피하려고 같은 개념의 용어를 바꾸지는 않는다.
|
|
21
|
+
4. **형식은 이해를 돕는 데 쓴다.** 비교·절차·기술 문서에 필요한 표·목록·제목은 살린다. 장식용 강조와 반복되는 문단 틀만 줄인다. 문장마다 줄을 바꾸거나, 모든 목록을 산문으로 만들거나, 문장 길이와 종결어미를 억지로 다양하게 만들지 않는다.
|
|
22
|
+
5. **관용구는 어울릴 때만 쓴다.** ‘첫발을 떼다’, ‘손이 많이 가다’처럼 익숙한 표현도 뜻과 상황에 맞을 때 쓴다. 더 쉽게 전달되지 않으면 평이하게 쓴다. 사람처럼 보이려고 감탄·반문·유행어·꾸며 낸 경험을 넣지 않는다.
|
|
23
|
+
6. **표현 하나만으로 문제 삼지 않는다.** ‘~를 통해’, ‘~할 수 있다’, 피동문·격식체·영어 자체는 수정 사유가 아니다. 문맥·반복·독자·목적을 기준으로 판단하고, 이미 자연스러운 문장은 그대로 둔다.
|
|
24
|
+
|
|
25
|
+
## 퇴고와 확인
|
|
26
|
+
|
|
27
|
+
의미 전달을 방해하는 구문부터 고치고, 반복과 군더더기는 그다음에 다듬는다. 표현·문장 단위로 고치되, 문제 해결에 필요하면 문장을 나누거나 문단을 조정한다. 단순한 취향 차이는 수정 사유로 삼지 않는다.
|
|
28
|
+
|
|
29
|
+
원문이 있으면 결과와 대조한다. 정보·주장·수치·주체·인과·조건·부정·가능성·의무·현재 상태가 달라지지 않았는지 확인한다. 요청 없이 새 주장·근거·감정·비유를 보태지 않는다. 사실 오류를 발견하면 문체 수정과 구분해 알린다.
|
|
30
|
+
|
|
31
|
+
고유명사·직접 인용·코드·명령어·경로·옵션값·링크 주소는 별도 요청 없이 바꾸지 않는다. 의미가 달라진 수정만 되돌리거나 바로잡고, 새로운 반복과 마크다운 깨짐이 없는지 확인한다.
|
|
32
|
+
|
|
33
|
+
## 예시
|
|
34
|
+
|
|
35
|
+
문맥에 따른 예시이며, 일괄 치환 규칙이 아니다.
|
|
36
|
+
|
|
37
|
+
- “skill과 rule을 적용한다.” → “스킬과 규칙을 적용한다.” (설명문)
|
|
38
|
+
- “해당 문제에 대한 검토를 수행합니다.” → “그 문제를 검토합니다.”
|
|
39
|
+
- “프로젝트의 첫 번째 발걸음을 취했다.” → “프로젝트의 첫발을 뗐다.”
|
|
40
|
+
- “예약이 취소되었습니다.” / “효과가 있을 수 있습니다.” → 그대로 둔다.
|
|
41
|
+
- “하네스로 작동한다.” → “하네스를 목표로 한다.”로 바꾸지 않는다. 현재 상태를 목표로 바꾸면 뜻이 달라진다.
|
|
42
|
+
|
|
43
|
+
## 출력
|
|
44
|
+
|
|
45
|
+
기본은 완성된 글만 제시한다. 진단을 요청하면 중요한 문제의 실제 표현·영향·수정 방향을 짧게 설명한다. 문제 개수나 진단표를 억지로 채우지 않는다. 진단만 요청하면 수정본은 만들지 않는다. 문제가 없으면 원문을 유지하고, 진단 요청에는 ‘큰 문제 없음’이라고 답한다.
|
|
@@ -129,8 +129,8 @@ Release 완성 작업 → Parent: Initiative Exit Criterion
|
|
|
129
129
|
|
|
130
130
|
- **audit-service-gaps** — *detects* north-star gaps end-to-end. This workflow consumes those gaps
|
|
131
131
|
as the evidence in step 2.
|
|
132
|
-
-
|
|
133
|
-
|
|
132
|
+
- The project's ADR + plan-SSOT conventions — the persistence mechanism (step 5) reuses them
|
|
133
|
+
rather than reinventing.
|
|
134
134
|
|
|
135
135
|
> Audit and gap skills answer "what's wrong now?". This skill answers "where do we go, and in what
|
|
136
136
|
> order?" — and makes the answer durable.
|
|
@@ -1,21 +1,21 @@
|
|
|
1
1
|
---
|
|
2
|
-
name:
|
|
2
|
+
name: objective-brief
|
|
3
3
|
description: >-
|
|
4
|
-
Rewrite
|
|
5
|
-
success_criteria, boundaries, autonomy, verification,
|
|
6
|
-
worker receives one judgeable definition of done instead
|
|
7
|
-
INBOUND reshapes a sprawling or half-formed request into that
|
|
8
|
-
context already on screen and
|
|
9
|
-
|
|
10
|
-
"프롬프트 구조화", and in English "turn this into a task brief",
|
|
11
|
-
|
|
12
|
-
worker, or a parallel lane
|
|
13
|
-
|
|
14
|
-
clarifying question is not a brief, and
|
|
15
|
-
|
|
4
|
+
Rewrite work about to be delegated, designed, or run over several steps into the canonical XML
|
|
5
|
+
brief — objective, inputs, invariants, success_criteria, boundaries, autonomy, verification,
|
|
6
|
+
communication, output_format — so the worker receives one judgeable definition of done instead
|
|
7
|
+
of prose. Runs in two directions: INBOUND reshapes a sprawling or half-formed request into that
|
|
8
|
+
shape, filling each field from context already on screen and dropping sections that do not
|
|
9
|
+
apply; OUTBOUND writes the prompt a spawned worker actually receives. Trigger on "브리프로 정리",
|
|
10
|
+
"작업 지시서로 만들어", "프롬프트 구조화", and in English "turn this into a task brief",
|
|
11
|
+
"structure this prompt". Fire unprompted before handing a multi-part task to a subagent, a
|
|
12
|
+
workflow worker, or a parallel lane, and before design or multi-step work at feature or project
|
|
13
|
+
scale. Do NOT fire on a one-line question, a lookup, a single edit, or a routine change — a
|
|
14
|
+
clarifying question is not a brief, and nine XML tags around a small ask cost more than they
|
|
15
|
+
buy.
|
|
16
16
|
---
|
|
17
17
|
|
|
18
|
-
#
|
|
18
|
+
# Objective Brief
|
|
19
19
|
|
|
20
20
|
A single shape for "here is the task". The point is not tidiness — it is that every field is a
|
|
21
21
|
place where an unstated assumption would otherwise stay unstated. A worker that receives prose
|
|
@@ -102,8 +102,9 @@ as a field that was considered and found empty, which is a claim you did not mak
|
|
|
102
102
|
|
|
103
103
|
## Inbound — normalize the request before acting on it
|
|
104
104
|
|
|
105
|
-
A request arrives as prose, a paste, or an ask that grew across
|
|
106
|
-
then work from the reshaped version.
|
|
105
|
+
A feature-scale or project-scale request arrives as prose, a paste, or an ask that grew across
|
|
106
|
+
three messages. Reshape it first, then work from the reshaped version. A one-line question, a
|
|
107
|
+
lookup, a single edit, or a routine change does not get a brief — answer it.
|
|
107
108
|
|
|
108
109
|
1. **Read the whole request before writing any field.** The objective is usually stated last.
|
|
109
110
|
2. **Fill fields from context you already have** — open files, the error text on screen, the
|
|
@@ -2,9 +2,9 @@
|
|
|
2
2
|
name: recurrence-prevention
|
|
3
3
|
description: >-
|
|
4
4
|
When the same defect, mistake, or incident happens AGAIN — a recurrence, not a one-off — verify
|
|
5
|
-
it against prior evidence (memory, rule
|
|
5
|
+
it against prior evidence (memory, rule 근거 links, git/CHANGELOG history), classify it as a
|
|
6
6
|
simple slip vs a complex harness problem, then escalate the countermeasure one level up the
|
|
7
|
-
ladder: record (1st) → forced rule with
|
|
7
|
+
ladder: record (1st) → forced rule with linked priors (2nd) → structural gate — test, hook, or
|
|
8
8
|
derive — once prose has failed (3rd+). Complex problems get countermeasure candidates designed
|
|
9
9
|
by a multi-persona panel instead of a quick patch. Use for "재발했어", "같은 실수 또 했네",
|
|
10
10
|
"이거 저번에도 그랬잖아", "재발방지 대책 등록해줘", "재발방지 룰 만들어", "this happened again",
|
|
@@ -34,7 +34,7 @@ is the next level up — not a louder version of the same level.**
|
|
|
34
34
|
- The user reports a recurrence or asks for a countermeasure: "재발했어", "같은 실수 또 했네",
|
|
35
35
|
"재발방지 대책 등록해줘", "postmortem this".
|
|
36
36
|
- You are fixing a defect and, while investigating, find a prior record of the same failure mode
|
|
37
|
-
(memory entry, rule
|
|
37
|
+
(memory entry, a rule's 신설 근거 link, CHANGELOG note) — even if nobody said "recurrence" out loud.
|
|
38
38
|
- A rule or gate that was supposed to prevent this class of failure existed **and was bypassed** —
|
|
39
39
|
that is itself a recurrence at the countermeasure level.
|
|
40
40
|
|
|
@@ -65,7 +65,8 @@ count basis — an unsourced generalization inflates counts and produces rule bl
|
|
|
65
65
|
same file does not. Write the signature down first — it decides everything after.
|
|
66
66
|
2. **Search prior evidence** for that signature, in order of reliability:
|
|
67
67
|
- durable memory (project memory entries, lessons/feedback notes)
|
|
68
|
-
- existing rule files and
|
|
68
|
+
- existing rule files and the occurrence links on their "신설 근거" line (a linked prior
|
|
69
|
+
occurrence = confirmed prior)
|
|
69
70
|
- `git log --grep`, CHANGELOG entries, ADRs, postmortem docs
|
|
70
71
|
- recorded observations, if the project keeps them — a session-observation digest that lists
|
|
71
72
|
commands and files repeated across sessions. Narrow but real: it turns "I kept redoing this"
|
|
@@ -114,8 +115,8 @@ one.)
|
|
|
114
115
|
|
|
115
116
|
| Level | When | Countermeasure | Characteristic failure of this level |
|
|
116
117
|
|---|---|---|---|
|
|
117
|
-
| **0 기록** | 1st occurrence | Fix + durable record: memory/lessons entry with **Why** it matters and **How to apply**, or
|
|
118
|
-
| **1 룰 강제 등록** | 2nd occurrence (the record failed) | Register a forced rule on the project's **always-loaded steering surface**, using the template below — one-line principle +
|
|
118
|
+
| **0 기록** | 1st occurrence | Fix + durable record: memory/lessons entry with **Why** it matters and **How to apply**, or an occurrence link on a related rule's 신설 근거 line if one already exists | Records don't steer — nothing re-reads them at the decision moment |
|
|
119
|
+
| **1 룰 강제 등록** | 2nd occurrence (the record failed) | Register a forced rule on the project's **always-loaded steering surface**, using the template below — one-line principle + a one-line "신설 근거" with links to each occurrence. `.claude/rules/<name>.md` (Claude Code) or a rules section in `AGENTS.md` (other CLIs) — and **verify it actually loads**: if the always-loaded context (CLAUDE.md / AGENTS.md) doesn't already pull that location in, reference the rule from it. A rule file nothing loads is still Level 0 with extra steps | Prose can be skimmed, forgotten under context pressure, or rationalized around |
|
|
119
120
|
| **2 구조적 게이트** | 3rd+ occurrence, **or** a registered countermeasure failed — bypassed *or* followed as designed yet insufficient | Deterministic enforcement that does not depend on the agent reading anything: a test gate that fails CI, a pre-action hook that blocks the command (where the CLI supports hooks), or **derive-to-single-source** so the drift is structurally impossible | Gates that never demonstrably fire; gates so noisy they get bypassed |
|
|
120
121
|
|
|
121
122
|
Load-bearing principle at Level 2: **comment warnings and doc reminders are not a blocking
|
|
@@ -165,12 +166,8 @@ a fact is free and always right.
|
|
|
165
166
|
```markdown
|
|
166
167
|
# <rule-name>
|
|
167
168
|
|
|
168
|
-
<One-line principle, stated as an imperative.>
|
|
169
|
-
|
|
170
|
-
| 사례 | 내용 |
|
|
171
|
-
|------|------|
|
|
172
|
-
| <date/version> | <what happened, one line> |
|
|
173
|
-
| <date/version> | <what happened, one line> |
|
|
169
|
+
<One-line principle, stated as an imperative.> **신설 근거: N회 재발** (<date/version> ·
|
|
170
|
+
<date/version>) — 발생별 기록은 <issue / ADR / postmortem link per occurrence>.
|
|
174
171
|
|
|
175
172
|
## 절대 원칙
|
|
176
173
|
|
|
@@ -180,12 +177,15 @@ a fact is free and always right.
|
|
|
180
177
|
## 위반 발견 시
|
|
181
178
|
|
|
182
179
|
1. 즉시 정정 보고 (무엇이 어떻게 위반이었는지 명시)
|
|
183
|
-
2.
|
|
180
|
+
2. 발생을 이슈·ADR 에 기록하고 durable memory 에 추가한 뒤, 위 근거 줄의 횟수와 링크를 올린다
|
|
184
181
|
3. 재위반이면 구조적 게이트(테스트/훅/derive)로 승격 — 프로즈는 이미 두 번 실패했다
|
|
185
182
|
```
|
|
186
183
|
|
|
187
|
-
The
|
|
188
|
-
|
|
184
|
+
The 근거 line is not decoration — it is the recurrence counter for the *next* occurrence, and it
|
|
185
|
+
is what makes the rule persuasive to a future agent deciding whether to comply. **Keep the count
|
|
186
|
+
and the links on the rule; keep the incident narratives in the linked issues/ADRs.** A standing
|
|
187
|
+
rule is re-read every session by every install, so each incident row it carries is a permanent
|
|
188
|
+
context cost — the link costs one line and still lets the next agent verify the count.
|
|
189
189
|
|
|
190
190
|
## Step 3b — Complex problem: multi-persona countermeasure design
|
|
191
191
|
|
|
@@ -231,7 +231,7 @@ would be a false ship. Before closing:
|
|
|
231
231
|
```
|
|
232
232
|
## 재발방지 보고
|
|
233
233
|
- Signature: <failure-mode class, one line>
|
|
234
|
-
- Count: N회 — evidence: <memory entry /
|
|
234
|
+
- Count: N회 — evidence: <memory entry / rule 근거 link / commit·CHANGELOG ref per occurrence>
|
|
235
235
|
- Classification: 단순 실수 | 복잡한 하네스 문제 (+ the discriminator that decided it)
|
|
236
236
|
- Countermeasure: Level 0 기록 | Level 1 룰 | Level 2 게이트 | 페르소나 설계 → <chosen option>
|
|
237
237
|
- Artifact: <path of memory entry / rule file / test or hook>
|