@uzysjung/agent-harness 26.149.0 → 26.151.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (65) hide show
  1. package/README.ko.md +1 -1
  2. package/README.md +1 -1
  3. package/dist/{chunk-YSW3OLH4.js → chunk-3QBHZUVB.js} +164 -66
  4. package/dist/chunk-3QBHZUVB.js.map +1 -0
  5. package/dist/index.js +397 -293
  6. package/dist/index.js.map +1 -1
  7. package/dist/trust-tier-drift.js +5 -1
  8. package/dist/trust-tier-drift.js.map +1 -1
  9. package/package.json +1 -1
  10. package/templates/CLAUDE.md +145 -164
  11. package/templates/agents/build-error-resolver.md +1 -1
  12. package/templates/agents/plan-checker.md +1 -1
  13. package/templates/agents/reviewer.md +4 -5
  14. package/templates/antigravity/AGENTS.md.template +3 -23
  15. package/templates/codex/AGENTS.md.template +5 -56
  16. package/templates/hooks/protect-files.sh +4 -0
  17. package/templates/hooks/session-start.sh +57 -3
  18. package/templates/opencode/AGENTS.md.template +4 -52
  19. package/templates/opencode/opencode.json.template +0 -8
  20. package/templates/rules/change-management.md +0 -1
  21. package/templates/rules/cli-development.md +1 -1
  22. package/templates/rules/doc-governance.md +2 -0
  23. package/templates/rules/git-policy.md +1 -1
  24. package/templates/rules/ship-checklist.md +3 -3
  25. package/templates/rules/test-policy.md +3 -8
  26. package/templates/settings.json +1 -16
  27. package/templates/skills/agent-introspection-debugging/SKILL.md +1 -1
  28. package/templates/skills/audit-harness-fit/README.md +113 -0
  29. package/templates/skills/audit-harness-fit/SKILL.md +64 -433
  30. package/templates/skills/audit-harness-fit/evals/scenarios.yaml +222 -0
  31. package/templates/skills/audit-harness-fit/references/apply.md +66 -0
  32. package/templates/skills/audit-harness-fit/references/audit.md +160 -0
  33. package/templates/skills/audit-harness-fit/references/populate.md +74 -0
  34. package/templates/skills/audit-harness-fit/references/verification.md +123 -0
  35. package/templates/skills/audit-service-gaps/SKILL.md +6 -7
  36. package/templates/skills/clear-korean-communication/SKILL.md +8 -13
  37. package/templates/skills/compaction-handoff/SKILL.md +29 -12
  38. package/templates/skills/external-model-consult/SKILL.md +13 -24
  39. package/templates/skills/model-orchestration/SKILL.md +18 -15
  40. package/templates/skills/natural-korean/SKILL.md +45 -0
  41. package/templates/skills/north-star/SKILL.md +4 -6
  42. package/templates/skills/north-star/references/roadmap-method.md +2 -2
  43. package/templates/skills/{task-brief → objective-brief}/SKILL.md +17 -18
  44. package/templates/skills/recurrence-prevention/SKILL.md +16 -16
  45. package/dist/chunk-YSW3OLH4.js.map +0 -1
  46. package/templates/agents/code-reviewer.md +0 -237
  47. package/templates/agents/security-reviewer.md +0 -108
  48. package/templates/hooks/task-brief-nudge.sh +0 -57
  49. package/templates/skills/audit-harness-fit/references/official-criteria.md +0 -367
  50. package/templates/skills/continuous-learning-v2/SKILL.md +0 -361
  51. package/templates/skills/continuous-learning-v2/agents/observer-loop.sh +0 -362
  52. package/templates/skills/continuous-learning-v2/agents/observer.md +0 -189
  53. package/templates/skills/continuous-learning-v2/agents/session-guardian.sh +0 -150
  54. package/templates/skills/continuous-learning-v2/agents/start-observer.sh +0 -252
  55. package/templates/skills/continuous-learning-v2/config.json +0 -8
  56. package/templates/skills/continuous-learning-v2/hooks/observe.sh +0 -585
  57. package/templates/skills/continuous-learning-v2/scripts/detect-project.sh +0 -322
  58. package/templates/skills/continuous-learning-v2/scripts/instinct-cli.py +0 -1956
  59. package/templates/skills/continuous-learning-v2/scripts/lib/homunculus-dir.sh +0 -31
  60. package/templates/skills/continuous-learning-v2/scripts/migrate-homunculus.sh +0 -68
  61. package/templates/skills/continuous-learning-v2/scripts/test_parse_instinct.py +0 -1420
  62. package/templates/skills/humanize-korean/SKILL.md +0 -228
  63. package/templates/skills/spec-scaling/SKILL.md +0 -89
  64. package/templates/skills/strategic-compact/SKILL.md +0 -145
  65. package/templates/skills/strategic-compact/suggest-compact.sh +0 -54
@@ -7,13 +7,12 @@ description: >-
7
7
  user-perspective (UX) — and enumerates concrete, severity-ranked gaps; BENCHMARK verifies how a
8
8
  reference service closes each one and PROPOSES a differentiated close. VERIFY, CHANGE-IMPACT,
9
9
  DRIFT and FULL extend the same loop to post-fix closure, baseline changes, doc-vs-code drift, and
10
- the whole-service sweep. Use when the user says any of: "북극성 기준으로 부족한 점", "사용자
11
- 관점에서 부족한 점", "다른 벤치마크 서비스는 이 부분을 어떻게 해결했는지", "갭분석", "레퍼런스
12
- 서비스랑 비교해서 부족한 점 찾아줘" — or the English equivalents: "gap analysis", "what are we
13
- missing vs the ideal/north-star", "benchmark against reference services", "audit this service".
14
- Fires for both Korean and English phrasing. Do NOT use it to *define* product direction (that is
15
- north-star), to review ONE standalone artifact's prose (that is multi-persona-review), or to turn
16
- an unverified benchmark claim into a fact.
10
+ the whole-service sweep. Use when the user says any of: "북극성 기준으로 부족한 점", "갭분석",
11
+ "다른 벤치마크 서비스는 이 부분을 어떻게 해결했는지", "레퍼런스 서비스랑 비교해서 부족한 점
12
+ 찾아줘" — or the English "gap analysis", "benchmark against reference services", "audit this
13
+ service". Do NOT use it to *define* product direction (that is north-star), to review ONE
14
+ standalone artifact's prose (that is multi-persona-review), or to turn an unverified benchmark
15
+ claim into a fact.
17
16
  ---
18
17
 
19
18
  # Audit Service Gaps (reverse + competitive)
@@ -4,21 +4,16 @@ description: >-
4
4
  Make a technical explanation land, and turn the decision at the end of it into
5
5
  something the reader can approve in one pass. Two halves of one job: (1) EXPLAIN
6
6
  — fix the referent first (one name often points at two things), lead with who is
7
- affected and what changes, put file paths and symbols after the claim as evidence;
7
+ affected and what changes, put file paths after the claim as evidence;
8
8
  (2) DECIDE — present approval/choice moments in the user's four-part format
9
- 전후맥락 (context) → 추천 + 이유 (recommendation) → UI/UX 형태 (a scannable
10
- table/option-list) → ASIS→TOBE contrast, led by the recommendation so the user can
9
+ 전후맥락 (context) → 추천 + 이유 (recommendation) → UI/UX 형태 (a scannable table) → ASIS→TOBE contrast, led by the recommendation so the user can
11
10
  say yes fast. Run it whenever you explain a bug, a cause, or what your change did
12
- — especially the moment the reader says they don't follow ("뭔 소리야", "쉽게
13
- 설명해줘", "이해가 안 돼", "I don't follow", "in plain terms"), or when your draft
14
- opens with a file path or symbol name — and whenever you are about to ask
15
- "should I do this?". Triggers on the user's verbatim phrases "ASIS TOBE로 설명",
16
- "ASIS-TOBE로 알려줘", "화면으로 ASIS TOBE로 설명", "의사결정 / 컨펌 요청",
17
- "이거 진행할까요?", the softer "다음 진행할 것들 알려줘", and the English
18
- equivalents "present this as ASIS/TOBE", "give me the as-is to-be", "should I do
19
- A or B", "ask for my approval", "lay out the options". Do NOT fire for pure
20
- information with no decision in it, for trivial reversible actions you would just
21
- do, or for context-free word/sentence translation — that is ordinary translation.
11
+ — especially the moment the reader says they don't follow ("뭔 소리야", "쉽게 설명해줘",
12
+ "이해가 안 돼", "I don't follow") — and whenever you are about to ask "should I do
13
+ this?" ("ASIS TOBE로 설명", "이거 진행할까요?", "should I do A or B"). Do NOT fire
14
+ for pure information with no decision in it, for trivial reversible actions you
15
+ would just do, or for context-free word/sentence translation — that is ordinary
16
+ translation.
22
17
  ---
23
18
 
24
19
  # Clear Korean Communication
@@ -22,7 +22,8 @@ This protocol is **snapshot-based, not append-based**:
22
22
  - Git and PR state: authoritative implementation snapshot.
23
23
  - `/compact` line: short pointer to the anchor, not another summary.
24
24
 
25
- > `strategic-compact` decides **when** to compact. This skill defines **how** to checkpoint.
25
+ > Claude Code compacts on its own when the context fills up. This skill defines **how** to
26
+ > checkpoint before that happens — or before you compact by hand.
26
27
 
27
28
  ## Goals
28
29
 
@@ -55,21 +56,38 @@ Do not run after compaction without first reconstructing the best available stat
55
56
  - Historical archives are disabled by default. If explicitly required, keep them under
56
57
  `.handoff/archive/`, apply a retention limit, and never auto-load them on resume.
57
58
 
58
- ### `MEMORY.md` — durable facts only
59
+ ### `MEMORY.md` — rules, not history
59
60
 
60
- Keep:
61
+ Memory is loaded **every session**, so every line is a standing cost. What earns that cost is
62
+ **principles, recurrence countermeasures, and facts you actually need to do the work** — not a
63
+ record of what was done.
61
64
 
62
- - stable purpose, scope, constraints, invariants, and operating principles;
63
- - durable repository or user preferences;
64
- - pointers to active ADRs and `.handoff/CURRENT.md`.
65
+ **Judge every entry — the ones you are adding AND the ones already there — with three questions:**
65
66
 
66
- Remove or exclude:
67
+ 1. **Does it change what I do next time?** If not, don't write it. "We shipped X" changes nothing.
68
+ 2. **Does it already live somewhere?** Rules, the project's instruction files, ADRs, skills, and
69
+ git history are each a source of truth. If the fact is there, that place owns it — do not keep
70
+ a copy here. **A duplicated fact is guaranteed to rot on one side**, and you cannot tell which.
71
+ 3. **Is it finished?** Completed cycles, release logs, and version history belong to the
72
+ changelog, ADRs, and git — not here.
67
73
 
68
- - current branch, test failure, task progress, temporary blocker, raw output;
69
- - completed session history, old anchors, duplicated repository content.
74
+ What survives all three: **operating principles · countermeasures for repeated mistakes ·
75
+ facts that cannot be derived from the repository** (another tool's flags, limits, and policies;
76
+ standing decisions such as "we accepted this risk, do not re-open it").
70
77
 
71
- Update or replace existing entries; do not append near-duplicates. Remove superseded facts. Target:
72
- **200 lines / 20 KB maximum**, unless the repository defines another limit.
78
+ **Re-judge the whole index at every handoff, not just the new lines.** An index only ever grows
79
+ unless something forces the question, and this is that moment.
80
+
81
+ Prefer updating an existing entry over adding a near-duplicate. To drop one, **move the file to
82
+ `archive/` rather than deleting it** — it leaves the index (so it stops loading) while staying
83
+ recoverable.
84
+
85
+ **Size is a symptom, not the standard.** Keep the index under **200 lines / 20 KB** (unless the
86
+ repository sets another limit), but being under it is *not* evidence the index is healthy: a short
87
+ index full of duplicates and finished history still fails all three questions. Measured case: an
88
+ index at 67 lines / 20.7 KB passed the size rule while **48 of its entries were dead** — completed
89
+ cycle records and copies of facts already owned by rules and ADRs. Re-judged against the three
90
+ questions, it came out at 23 lines / 6.0 KB.
73
91
 
74
92
  ### `docs/decisions/` — durable decisions only
75
93
 
@@ -308,5 +326,4 @@ a minimal manual resume anchor in the response.
308
326
 
309
327
  ## Related skills
310
328
 
311
- - **strategic-compact** — decides when to compact.
312
329
  - **git-policy Session Cleanup** — defines repository and PR cleanup expectations.
@@ -1,30 +1,19 @@
1
1
  ---
2
2
  name: external-model-consult
3
3
  description: >-
4
- Consult a second, non-Claude model through a bundled wrapper for the four things
5
- an external round-trip actually buys: (1) natural, native-sounding KOREAN phrasing
6
- via Google Gemini (Antigravity `agy` CLI) — copy, UI microcopy, marketing/brochure
7
- text, toasts, user-facing messages, translations, rewrites; (2) a MULTI-PERSONA /
8
- second-opinion review of a design, plan, spec, PR, or piece of writing; (3) CONCISE,
9
- well-STRUCTURED writing via OpenAI Codex (`codex exec`) — tightening verbose prose,
10
- restructuring a doc into a clean outline / tables / sections, executive summaries,
11
- README skeletons, changelog entries; and (4) IMAGE GENERATION as real PNG/JPG files
12
- on disk (labeled flowcharts / architecture / sequence diagrams are NOT this — render
13
- those natively as Mermaid). Use whenever Korean text needs to read naturally rather
14
- than translated, whenever the user says the Korean "sounds awkward / 어색해 /
15
- 자연스럽게 다듬어줘", whenever you are about to hand-write polished Korean copy
16
- yourself, whenever a document needs to get SHORTER and better ORGANIZED (not
17
- prettier-sounding), whenever you want an independent non-Claude critique, or when
18
- the user asks for a generated image. Triggers on "gemini 한테 물어봐 / gemini 로
19
- 다듬어 / gemini로 이미지 만들어줘 / nano banana / agy / antigravity / 제3자 관점 /
20
- second opinion / codex한테 물어봐 / codex로 정리해 / 간결하게 정리해줘 / 구조화해줘 /
21
- 문서 구조 잡아줘 / 이미지 만들어줘 / 그림 생성해줘", and in English "ask codex",
22
- "ask gemini", "tighten this up", "make this concise", "restructure this doc",
23
- "generate an image". Returns candidates/files for the user to choose from; never
24
- auto-applies. Do NOT use for deterministic transforms (rename, reformat, sort),
25
- labeled diagrams, internal logs or code identifiers, anything needing repo secrets,
26
- a native parallel-subagent panel (that is `multi-persona-review` where installed),
27
- or when the user explicitly wants YOUR answer.
4
+ Consult a second, non-Claude model through a bundled wrapper for four things:
5
+ (1) natural, native-sounding KOREAN phrasing via Google Gemini — copy, UI microcopy,
6
+ marketing text; (2) a MULTI-PERSONA / second-opinion review of a design, plan or spec;
7
+ (3) CONCISE, well-STRUCTURED writing via OpenAI Codex — tightening prose, restructuring
8
+ a doc; and (4) IMAGE GENERATION as real image files on disk (not labeled diagrams —
9
+ use Mermaid). Fire it when Korean reads translated, or when you would otherwise
10
+ hand-write polished Korean yourself. Triggers on "어색해", "자연스럽게 다듬어줘",
11
+ "gemini 한테 물어봐", "nano banana", "제3자 관점", "codex한테 물어봐",
12
+ "간결하게 정리해줘", "이미지 만들어줘", and in English "ask gemini", "ask codex",
13
+ "generate an image". Returns candidates to choose from; never auto-applies. Do NOT use
14
+ for deterministic transforms (rename, reformat, sort), labeled diagrams, internal logs
15
+ or identifiers, anything needing repo secrets, a native subagent panel
16
+ (`multi-persona-review`), or when the user explicitly wants YOUR answer.
28
17
  ---
29
18
 
30
19
  # external-model-consult
@@ -2,25 +2,28 @@
2
2
  name: model-orchestration
3
3
  description: >-
4
4
  Apply the fixed model-role and thinking-effort policy whenever work is delegated to subagents
5
- or a model/effort choice is made: the orchestrator (top-tier model, Fable) DIRECTLY owns 설계·
6
- 기획·분배·리뷰 — sets service direction, arbitrates and reviews plan/spec documents (with
7
- multi-persona-review), improves shipped features, and hunts performance/security problems;
8
- core implementation, test authoring/execution (E2E included), and code verification/V&V go to
9
- Opus at xhigh or above; repetitive/simple implementation (applying an established pattern)
10
- goes to Sonnet at high or above — Sonnet is never used for tests or verification;
11
- plan/spec drafts may be produced by Opus from a Fable direction brief, but the
12
- review/decision is always Fable's. Never delegate below the effort floors. Use whenever you
13
- are about to spawn an Agent/Task/Workflow worker, pick a model for a subtask, set a
14
- thinking/effort level, assign verification, or hand off orchestration because the current
15
- model's quota is exhausted. Trigger on "위임해", "에이전트로 돌려", "오케스트레이션",
16
- "모델 역할분담", "어떤 모델로", "effort 얼마로", "thinking level", "서브에이전트", or in English
17
- "delegate this", "spawn an agent for", "which model should", "route this task", "verify with",
18
- "orchestrate". Fire even when the user doesn't name the policy — any delegation decision is
19
- in scope.
5
+ or a model/effort choice is made: the orchestrator (top-tier model, Fable) DIRECTLY owns
6
+ 설계·기획·분배·리뷰 — sets direction, reviews plan/spec docs, improves shipped features, and hunts
7
+ perf/security problems; core implementation, test authoring/execution
8
+ (E2E included), and code verification/V&V go to Opus at xhigh or above;
9
+ repetitive/simple implementation goes to Sonnet at high or above — Sonnet is never used for
10
+ tests or verification; plan/spec drafts may be produced by Opus from a Fable direction brief,
11
+ but the review/decision is always Fable's. Never delegate below the effort floors. Use whenever
12
+ you are about to spawn an Agent/Task/Workflow worker, pick a model for a subtask, set a
13
+ thinking/effort level, or assign verification. Trigger on "위임해", "오케스트레이션", "모델 역할분담", "어떤 모델로",
14
+ "effort 얼마로", "서브에이전트", "delegate this", "which model should". Fire even when the
15
+ user doesn't name the policy — any delegation is in scope.
20
16
  ---
21
17
 
22
18
  # Model Orchestration Policy
23
19
 
20
+ **What is policy and what is vocabulary.** The policy is the role split — orchestrator ≠ builder ≠
21
+ verifier — and the rule that verification never goes to a lower tier than the work it verifies.
22
+ `Fable` · `Opus` · `Sonnet` and the effort levels (`xhigh` · `high` · `max`) are Claude Code
23
+ vocabulary. On Codex, OpenCode, or Antigravity apply the same split with that vendor's top tier for
24
+ orchestration and verification and its mid tier for repetitive implementation; an effort knob that
25
+ does not exist there is not a violation — putting verification on a lower tier is.
26
+
24
27
  A fixed role split between model tiers, set by the user (2026-07-04, revised 2026-07-07 and
25
28
  2026-08-02 — *"설계·분배·기획·리뷰는 Fable5, 핵심 구현·테스트·검증은 opus5, 반복·단순 구현은
26
29
  sonnet"*). The premise is
@@ -0,0 +1,45 @@
1
+ ---
2
+ name: natural-korean
3
+ description: 한국어 답변·글쓰기·번역·퇴고에 사용한다. 뜻과 말투를 살리면서 번역투·불필요한 영어·상투구·기계적 반복을 줄인다. "자연스럽게 다듬어줘", "AI 티 없애줘", "번역투 고쳐줘", "문체 진단" 요청에도 적용한다.
4
+ ---
5
+
6
+ # 자연스러운 한국어
7
+
8
+ 쉽고 익숙한 한국어로 뜻을 정확하게 전달한다. 사람처럼 보이게 꾸미기보다, 실제로 어색한 표현을 고친다.
9
+
10
+ ## 작업 방식
11
+
12
+ - **새 글·답변:** 독자와 목적에 맞게 쓴다. 설명과 답변은 핵심부터 말하고, 별도 요청이 없는 대화는 자연스러운 존댓말을 쓴다.
13
+ - **번역:** 단어보다 뜻을 옮긴다. 정보와 뉘앙스는 유지하되, 어순·문장 구성·관용구는 자연스러운 한국어로 바꾼다.
14
+ - **퇴고:** 글 전체의 문맥을 파악한 뒤, 실제 문제를 진단하고 필요한 부분만 고친다. 별도 요청이 없으면 원래 말투·구조를 유지한다. 재작성이나 말투 변경을 요청하면 그 범위에서 바꾼다.
15
+
16
+ ## 표현 원칙
17
+
18
+ 1. **독자에게 익숙한 말을 쓴다.** 전문용어의 뜻은 보존하되, 설명문에서는 ‘스킬·규칙·프롬프트’처럼 익숙한 표기를 우선한다. 제품명·식별자·통용 약어는 유지한다. 낯선 용어는 필요할 때 짧게 설명하고, 원어 병기는 구분에 필요한 경우만 한다. 뜻이 불분명한 용어를 짐작으로 풀지 않는다.
19
+ 2. **누가 무엇을 하는지 분명히 쓴다.** 명사를 길게 잇기보다 동사로 직접 표현한다. 불필요한 대명사·명사화·간접 표현을 덜되, 주체와 논리관계는 남긴다. 긴 문장은 필요할 때 나누고, 짧게 쓰려고 조사·어미·필요한 설명을 빼지 않는다.
20
+ 3. **반복과 군더더기를 줄인다.** 내용에 보탬이 없는 칭찬·예고·과장·결론 반복은 덜어낸다. 접속사·완곡 표현·같은 문형도 문맥상 불필요하거나 반복이 거슬릴 때만 고친다. 반복을 피하려고 같은 개념의 용어를 바꾸지는 않는다.
21
+ 4. **형식은 이해를 돕는 데 쓴다.** 비교·절차·기술 문서에 필요한 표·목록·제목은 살린다. 장식용 강조와 반복되는 문단 틀만 줄인다. 문장마다 줄을 바꾸거나, 모든 목록을 산문으로 만들거나, 문장 길이와 종결어미를 억지로 다양하게 만들지 않는다.
22
+ 5. **관용구는 어울릴 때만 쓴다.** ‘첫발을 떼다’, ‘손이 많이 가다’처럼 익숙한 표현도 뜻과 상황에 맞을 때 쓴다. 더 쉽게 전달되지 않으면 평이하게 쓴다. 사람처럼 보이려고 감탄·반문·유행어·꾸며 낸 경험을 넣지 않는다.
23
+ 6. **표현 하나만으로 문제 삼지 않는다.** ‘~를 통해’, ‘~할 수 있다’, 피동문·격식체·영어 자체는 수정 사유가 아니다. 문맥·반복·독자·목적을 기준으로 판단하고, 이미 자연스러운 문장은 그대로 둔다.
24
+
25
+ ## 퇴고와 확인
26
+
27
+ 의미 전달을 방해하는 구문부터 고치고, 반복과 군더더기는 그다음에 다듬는다. 표현·문장 단위로 고치되, 문제 해결에 필요하면 문장을 나누거나 문단을 조정한다. 단순한 취향 차이는 수정 사유로 삼지 않는다.
28
+
29
+ 원문이 있으면 결과와 대조한다. 정보·주장·수치·주체·인과·조건·부정·가능성·의무·현재 상태가 달라지지 않았는지 확인한다. 요청 없이 새 주장·근거·감정·비유를 보태지 않는다. 사실 오류를 발견하면 문체 수정과 구분해 알린다.
30
+
31
+ 고유명사·직접 인용·코드·명령어·경로·옵션값·링크 주소는 별도 요청 없이 바꾸지 않는다. 의미가 달라진 수정만 되돌리거나 바로잡고, 새로운 반복과 마크다운 깨짐이 없는지 확인한다.
32
+
33
+ ## 예시
34
+
35
+ 문맥에 따른 예시이며, 일괄 치환 규칙이 아니다.
36
+
37
+ - “skill과 rule을 적용한다.” → “스킬과 규칙을 적용한다.” (설명문)
38
+ - “해당 문제에 대한 검토를 수행합니다.” → “그 문제를 검토합니다.”
39
+ - “프로젝트의 첫 번째 발걸음을 취했다.” → “프로젝트의 첫발을 뗐다.”
40
+ - “예약이 취소되었습니다.” / “효과가 있을 수 있습니다.” → 그대로 둔다.
41
+ - “하네스로 작동한다.” → “하네스를 목표로 한다.”로 바꾸지 않는다. 현재 상태를 목표로 바꾸면 뜻이 달라진다.
42
+
43
+ ## 출력
44
+
45
+ 기본은 완성된 글만 제시한다. 진단을 요청하면 중요한 문제의 실제 표현·영향·수정 방향을 짧게 설명한다. 문제 개수나 진단표를 억지로 채우지 않는다. 진단만 요청하면 수정본은 만들지 않는다. 문제가 없으면 원문을 유지하고, 진단 요청에는 ‘큰 문제 없음’이라고 답한다.
@@ -10,12 +10,10 @@ description: >-
10
10
  should go next. Sits one layer above SPEC/PRD — answers 'why and where to', not
11
11
  'what and how'. Fires on the user's real phrasings: "앞으로 어떤 방향으로
12
12
  개선·발전시킬지 고민해봐", "NORTH.md / NORTH_STAR 보고 나아갈 방향 + 기능 제안",
13
- "나아갈 방향 + 기능제안 (수용 → 계획 수립하고 메모리에 기록)", "북극성 정렬 로드맵",
14
- as well as the English equivalents "what direction should we take next", "propose
15
- a roadmap / feature backlog from the north star", "plan the next milestones and
16
- save it to memory". Do NOT use it to find what is broken right now — detecting
17
- bugs, gaps, or quality regressions belongs to the audit/gap skills; this skill
18
- consumes their findings and DIRECTS forward planning.
13
+ "북극성 정렬 로드맵", and the English "what direction should we take next",
14
+ "propose a roadmap from the north star". Do NOT use it to find what is broken
15
+ right now — detecting bugs, gaps, or quality regressions belongs to the audit/gap
16
+ skills; this skill consumes their findings and DIRECTS forward planning.
19
17
  ---
20
18
 
21
19
  # North Star
@@ -129,8 +129,8 @@ Release 완성 작업 → Parent: Initiative Exit Criterion
129
129
 
130
130
  - **audit-service-gaps** — *detects* north-star gaps end-to-end. This workflow consumes those gaps
131
131
  as the evidence in step 2.
132
- - **strategic-compact** / the project's ADR + plan-SSOT conventions — the persistence mechanism
133
- (step 5) reuses them rather than reinventing.
132
+ - The project's ADR + plan-SSOT conventions — the persistence mechanism (step 5) reuses them
133
+ rather than reinventing.
134
134
 
135
135
  > Audit and gap skills answer "what's wrong now?". This skill answers "where do we go, and in what
136
136
  > order?" — and makes the answer durable.
@@ -1,23 +1,21 @@
1
1
  ---
2
- name: task-brief
2
+ name: objective-brief
3
3
  description: >-
4
- Rewrite a task request into the canonical XML brief — objective, inputs, invariants,
5
- success_criteria, boundaries, autonomy, verification, communication, output_format — so the
6
- worker receives one judgeable definition of done instead of prose. Runs in two directions:
7
- INBOUND reshapes a sprawling or half-formed request (pasted requirements, a wall of background,
8
- a request that grew across several messages) into that shape, filling each field from context
9
- already on screen and deleting sections that do not apply; OUTBOUND writes the prompt that a
10
- spawned worker actually receives, so nobody hand-rolls a one-off prompt shape per delegation.
11
- Trigger on "브리프로 정리", "작업 지시서로 만들어", "프롬프트 구조화", "브리프 만들어줘",
12
- "이 요청 정리해줘", and in English "turn this into a task brief", "structure this prompt",
13
- "write the brief for this", "draft the spawn prompt". Fire unprompted the moment you are about
14
- to hand a multi-part task to a subagent, a workflow worker, or a parallel lane. Do NOT fire on a
15
- one-line question, a lookup, or an ordinary conversational exchange where you simply need one
16
- more piece of information — asking a clarifying question is not a brief, and wrapping a
17
- one-sentence ask in nine XML tags costs more than it buys.
4
+ Rewrite work about to be delegated, designed, or run over several steps into the canonical XML
5
+ brief — objective, inputs, invariants, success_criteria, boundaries, autonomy, verification,
6
+ communication, output_format — so the worker receives one judgeable definition of done instead
7
+ of prose. Runs in two directions: INBOUND reshapes a sprawling or half-formed request into that
8
+ shape, filling each field from context already on screen and dropping sections that do not
9
+ apply; OUTBOUND writes the prompt a spawned worker actually receives. Trigger on "브리프로 정리",
10
+ "작업 지시서로 만들어", "프롬프트 구조화", and in English "turn this into a task brief",
11
+ "structure this prompt". Fire unprompted before handing a multi-part task to a subagent, a
12
+ workflow worker, or a parallel lane, and before design or multi-step work at feature or project
13
+ scale. Do NOT fire on a one-line question, a lookup, a single edit, or a routine change — a
14
+ clarifying question is not a brief, and nine XML tags around a small ask cost more than they
15
+ buy.
18
16
  ---
19
17
 
20
- # Task Brief
18
+ # Objective Brief
21
19
 
22
20
  A single shape for "here is the task". The point is not tidiness — it is that every field is a
23
21
  place where an unstated assumption would otherwise stay unstated. A worker that receives prose
@@ -104,8 +102,9 @@ as a field that was considered and found empty, which is a claim you did not mak
104
102
 
105
103
  ## Inbound — normalize the request before acting on it
106
104
 
107
- A request arrives as prose, a paste, or an ask that grew across three messages. Reshape it first,
108
- then work from the reshaped version.
105
+ A feature-scale or project-scale request arrives as prose, a paste, or an ask that grew across
106
+ three messages. Reshape it first, then work from the reshaped version. A one-line question, a
107
+ lookup, a single edit, or a routine change does not get a brief — answer it.
109
108
 
110
109
  1. **Read the whole request before writing any field.** The objective is usually stated last.
111
110
  2. **Fill fields from context you already have** — open files, the error text on screen, the
@@ -2,9 +2,9 @@
2
2
  name: recurrence-prevention
3
3
  description: >-
4
4
  When the same defect, mistake, or incident happens AGAIN — a recurrence, not a one-off — verify
5
- it against prior evidence (memory, rule case tables, git/CHANGELOG history), classify it as a
5
+ it against prior evidence (memory, rule 근거 links, git/CHANGELOG history), classify it as a
6
6
  simple slip vs a complex harness problem, then escalate the countermeasure one level up the
7
- ladder: record (1st) → forced rule with a case table (2nd) → structural gate — test, hook, or
7
+ ladder: record (1st) → forced rule with linked priors (2nd) → structural gate — test, hook, or
8
8
  derive — once prose has failed (3rd+). Complex problems get countermeasure candidates designed
9
9
  by a multi-persona panel instead of a quick patch. Use for "재발했어", "같은 실수 또 했네",
10
10
  "이거 저번에도 그랬잖아", "재발방지 대책 등록해줘", "재발방지 룰 만들어", "this happened again",
@@ -34,7 +34,7 @@ is the next level up — not a louder version of the same level.**
34
34
  - The user reports a recurrence or asks for a countermeasure: "재발했어", "같은 실수 또 했네",
35
35
  "재발방지 대책 등록해줘", "postmortem this".
36
36
  - You are fixing a defect and, while investigating, find a prior record of the same failure mode
37
- (memory entry, rule case table, CHANGELOG note) — even if nobody said "recurrence" out loud.
37
+ (memory entry, a rule's 신설 근거 link, CHANGELOG note) — even if nobody said "recurrence" out loud.
38
38
  - A rule or gate that was supposed to prevent this class of failure existed **and was bypassed** —
39
39
  that is itself a recurrence at the countermeasure level.
40
40
 
@@ -65,7 +65,8 @@ count basis — an unsourced generalization inflates counts and produces rule bl
65
65
  same file does not. Write the signature down first — it decides everything after.
66
66
  2. **Search prior evidence** for that signature, in order of reliability:
67
67
  - durable memory (project memory entries, lessons/feedback notes)
68
- - existing rule files and their case tables (a matching case-table row = confirmed prior)
68
+ - existing rule files and the occurrence links on their "신설 근거" line (a linked prior
69
+ occurrence = confirmed prior)
69
70
  - `git log --grep`, CHANGELOG entries, ADRs, postmortem docs
70
71
  - recorded observations, if the project keeps them — a session-observation digest that lists
71
72
  commands and files repeated across sessions. Narrow but real: it turns "I kept redoing this"
@@ -114,8 +115,8 @@ one.)
114
115
 
115
116
  | Level | When | Countermeasure | Characteristic failure of this level |
116
117
  |---|---|---|---|
117
- | **0 기록** | 1st occurrence | Fix + durable record: memory/lessons entry with **Why** it matters and **How to apply**, or a case-table row if a related rule already exists | Records don't steer — nothing re-reads them at the decision moment |
118
- | **1 룰 강제 등록** | 2nd occurrence (the record failed) | Register a forced rule on the project's **always-loaded steering surface**, using the template below — one-line principle + case table. `.claude/rules/<name>.md` (Claude Code) or a rules section in `AGENTS.md` (other CLIs) — and **verify it actually loads**: if the always-loaded context (CLAUDE.md / AGENTS.md) doesn't already pull that location in, reference the rule from it. A rule file nothing loads is still Level 0 with extra steps | Prose can be skimmed, forgotten under context pressure, or rationalized around |
118
+ | **0 기록** | 1st occurrence | Fix + durable record: memory/lessons entry with **Why** it matters and **How to apply**, or an occurrence link on a related rule's 신설 근거 line if one already exists | Records don't steer — nothing re-reads them at the decision moment |
119
+ | **1 룰 강제 등록** | 2nd occurrence (the record failed) | Register a forced rule on the project's **always-loaded steering surface**, using the template below — one-line principle + a one-line "신설 근거" with links to each occurrence. `.claude/rules/<name>.md` (Claude Code) or a rules section in `AGENTS.md` (other CLIs) — and **verify it actually loads**: if the always-loaded context (CLAUDE.md / AGENTS.md) doesn't already pull that location in, reference the rule from it. A rule file nothing loads is still Level 0 with extra steps | Prose can be skimmed, forgotten under context pressure, or rationalized around |
119
120
  | **2 구조적 게이트** | 3rd+ occurrence, **or** a registered countermeasure failed — bypassed *or* followed as designed yet insufficient | Deterministic enforcement that does not depend on the agent reading anything: a test gate that fails CI, a pre-action hook that blocks the command (where the CLI supports hooks), or **derive-to-single-source** so the drift is structurally impossible | Gates that never demonstrably fire; gates so noisy they get bypassed |
120
121
 
121
122
  Load-bearing principle at Level 2: **comment warnings and doc reminders are not a blocking
@@ -165,12 +166,8 @@ a fact is free and always right.
165
166
  ```markdown
166
167
  # <rule-name>
167
168
 
168
- <One-line principle, stated as an imperative.> 신설 근거: N회 재발 (YYYY-MM-DD):
169
-
170
- | 사례 | 내용 |
171
- |------|------|
172
- | <date/version> | <what happened, one line> |
173
- | <date/version> | <what happened, one line> |
169
+ <One-line principle, stated as an imperative.> **신설 근거: N회 재발** (<date/version> ·
170
+ <date/version>) — 발생별 기록은 <issue / ADR / postmortem link per occurrence>.
174
171
 
175
172
  ## 절대 원칙
176
173
 
@@ -180,12 +177,15 @@ a fact is free and always right.
180
177
  ## 위반 발견 시
181
178
 
182
179
  1. 즉시 정정 보고 (무엇이 어떻게 위반이었는지 명시)
183
- 2. 본 사례표 + durable memory 에 추가
180
+ 2. 발생을 이슈·ADR 에 기록하고 durable memory 에 추가한 뒤, 위 근거 줄의 횟수와 링크를 올린다
184
181
  3. 재위반이면 구조적 게이트(테스트/훅/derive)로 승격 — 프로즈는 이미 두 번 실패했다
185
182
  ```
186
183
 
187
- The case table is not decoration — it is the recurrence counter for the *next* occurrence, and
188
- it is what makes the rule persuasive to a future agent deciding whether to comply.
184
+ The 근거 line is not decoration — it is the recurrence counter for the *next* occurrence, and it
185
+ is what makes the rule persuasive to a future agent deciding whether to comply. **Keep the count
186
+ and the links on the rule; keep the incident narratives in the linked issues/ADRs.** A standing
187
+ rule is re-read every session by every install, so each incident row it carries is a permanent
188
+ context cost — the link costs one line and still lets the next agent verify the count.
189
189
 
190
190
  ## Step 3b — Complex problem: multi-persona countermeasure design
191
191
 
@@ -231,7 +231,7 @@ would be a false ship. Before closing:
231
231
  ```
232
232
  ## 재발방지 보고
233
233
  - Signature: <failure-mode class, one line>
234
- - Count: N회 — evidence: <memory entry / case-table row / commit·CHANGELOG ref per occurrence>
234
+ - Count: N회 — evidence: <memory entry / rule 근거 link / commit·CHANGELOG ref per occurrence>
235
235
  - Classification: 단순 실수 | 복잡한 하네스 문제 (+ the discriminator that decided it)
236
236
  - Countermeasure: Level 0 기록 | Level 1 룰 | Level 2 게이트 | 페르소나 설계 → <chosen option>
237
237
  - Artifact: <path of memory entry / rule file / test or hook>