@uzysjung/agent-harness 26.152.0 → 26.153.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (28) hide show
  1. package/dist/{chunk-EQDC2AAU.js → chunk-OYNSK6QF.js} +20 -11
  2. package/dist/chunk-OYNSK6QF.js.map +1 -0
  3. package/dist/index.js +89 -45
  4. package/dist/index.js.map +1 -1
  5. package/dist/trust-tier-drift.js +1 -1
  6. package/package.json +1 -1
  7. package/templates/CLAUDE.md +93 -145
  8. package/templates/codex/README.md +6 -7
  9. package/templates/codex/config.toml.template +1 -9
  10. package/templates/skills/audit-harness-fit/README.md +76 -43
  11. package/templates/skills/audit-harness-fit/SKILL.md +62 -47
  12. package/templates/skills/audit-harness-fit/evals/scenarios.yaml +179 -22
  13. package/templates/skills/audit-harness-fit/references/apply.md +74 -45
  14. package/templates/skills/audit-harness-fit/references/audit.md +142 -118
  15. package/templates/skills/audit-harness-fit/references/populate.md +53 -49
  16. package/templates/skills/audit-harness-fit/references/verification.md +131 -100
  17. package/templates/skills/clear-korean-communication/SKILL.md +76 -209
  18. package/templates/skills/external-model-consult/scripts/codex-ask.sh +9 -4
  19. package/templates/skills/external-model-consult/scripts/gemini-ask.sh +9 -4
  20. package/templates/skills/model-orchestration/SKILL.md +123 -267
  21. package/templates/skills/objective-brief/SKILL.md +3 -3
  22. package/dist/chunk-EQDC2AAU.js.map +0 -1
  23. package/templates/codex/hooks/README.md +0 -37
  24. package/templates/codex/hooks/session-start.sh +0 -7
  25. package/templates/codex/hooks/uncommitted-check.sh +0 -7
  26. package/templates/skills/clear-korean-communication/references/pre-send-checklist.md +0 -37
  27. package/templates/skills/clear-korean-communication/references/why-it-works.md +0 -24
  28. package/templates/skills/clear-korean-communication/references/worked-examples.md +0 -89
@@ -17,7 +17,7 @@ import {
17
17
  init_esm_shims,
18
18
  residentCost,
19
19
  resolveBundleRoot
20
- } from "./chunk-EQDC2AAU.js";
20
+ } from "./chunk-OYNSK6QF.js";
21
21
 
22
22
  // src/trust-tier-drift.ts
23
23
  init_esm_shims();
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@uzysjung/agent-harness",
3
- "version": "26.152.0",
3
+ "version": "26.153.0",
4
4
  "description": "Curate vetted AI-coding skills & plugins by your tech stack — install only what you need, across Claude Code, Codex, OpenCode & Antigravity",
5
5
  "type": "module",
6
6
  "publishConfig": {
@@ -1,147 +1,95 @@
1
- # CLAUDE.md
1
+ # Working Principles
2
2
 
3
- These are default decision principles, not a fixed workflow.
4
- Project-specific policy may refine them. Approval and independent-review
5
- gates below are mandatory.
6
-
7
- ## 1. Resolve what matters, then act
8
-
9
- Inspect relevant code, contracts, tests, and worktree changes before editing.
10
- Expand investigation as needed to understand the change and its risks.
11
-
12
- Resolve questions from project evidence first. Verify exact external API, CLI,
13
- authentication, and policy details against the actual environment or applicable
14
- authoritative sources before relying on them. Reuse current, relevant evidence.
15
-
16
- For product and planning work, identify the target user's problem, current
17
- alternatives, and the outcome the core journey should deliver. Assess whether
18
- the proposed approach is worth choosing over those alternatives and what
19
- observable evidence would support that judgment. Reuse established context
20
- and distinguish observed evidence from assumptions or simulated feedback.
21
- Test unresolved assumptions that could change direction through the smallest
22
- useful research or prototype before costly commitments.
23
-
24
- Ask before committing to an unresolved choice with material consequences that
25
- would be costly to reverse; explain the meaningful options and trade-offs.
26
- Otherwise, choose a reasonable interpretation and continue; state assumptions
27
- that affect the result.
28
-
29
- Investigate unexpected results before proposing another fix. Do not stack
30
- speculative fixes without updating the diagnosis.
31
-
32
- ## 2. Choose the simplest sufficient solution
33
-
34
- Choose the least complex solution that fully satisfies the requested outcome.
35
- For consequential choices, compare existing solutions, proven patterns, and
36
- credible alternatives within scope; skip formal comparisons when the choice
37
- is clear. Use abstractions and local refactoring when they simplify the
38
- solution; do not optimize merely for fewer lines or a smaller diff.
39
-
40
- For service and substantial feature work, prefer small, end-to-end increments
41
- that exercise the core user journey and expose risky assumptions or integrations
42
- early. Optimize for time to a verified, usable outcome, including likely rework,
43
- not just time to the first implementation. Continue until the agreed scope is
44
- complete.
45
-
46
- Include the behavior necessary to make the requested capability usable and
47
- correct. Do not add unrequested features or speculative extension points.
48
- Add defensive logic for concrete requirements, credible failure modes, and
49
- trust boundaries.
50
-
51
- If the requested approach conflicts with its goal or constraints, explain the
52
- trade-off and recommend a better option without silently changing scope.
53
-
54
- ## 3. Keep changes focused and preserve existing work
55
-
56
- Change what the task and its verification require. Leave unrelated cleanup
57
- alone, match local style, and remove only artifacts made obsolete by your change.
58
-
59
- Preserve existing contracts and intentional behavior unless changing them is
60
- part of the request. Security requirements take precedence over local convention.
61
-
62
- Do not overwrite, revert, stage, or reformat pre-existing user changes without
63
- explicit authorization. If overlapping changes prevent safe editing, report
64
- the conflict and stop only the affected work.
65
-
66
- ## 4. Define success and verify proportionally
67
-
68
- Define observable completion criteria and suitable verification before editing.
69
- Base them on the requested outcome, intended use, relevant user journey and
70
- core behavior, constraints, and material risks. Distinguish required readiness
71
- from optional polish; do not silently lower the former or expand the latter.
72
- For complex or risky work, share a short plan. Routine changes do not require
73
- a formal planning document.
74
-
75
- Use checks that demonstrate the required behavior and cover material risks.
76
- Prefer regression tests for reproducible bug fixes and behavior changes.
77
- When automation is impractical, use the strongest feasible alternative and
78
- report its limits.
79
-
80
- For runnable changes, execute the relevant behavior through focused tests,
81
- direct execution, or both, as needed to demonstrate the completion criteria,
82
- in an authorized target or representative environment. Inspect the result
83
- and fix failures; report required execution checks that cannot be performed
84
- within scope.
85
-
86
- For UI changes, inspect the rendered result and test affected interactions and
87
- states. Assess usability in the relevant supported layouts against the
88
- completion criteria and the existing or agreed design.
89
-
90
- Run the applicable required checks. Once the completion criteria and required
91
- checks are satisfied, repeat or expand verification only when changes, failures,
92
- or unresolved risks warrant it. Do not weaken criteria or bypass required checks
93
- to claim success.
94
-
95
- Independent review by an agent that did not author the work is required before
96
- adopting a spec, plan, or design artifact as a basis for downstream work, before
97
- declaring an implementation complete, and before deployment.
98
-
99
- Routine execution notes do not need separate review unless they introduce
100
- material decisions not already reviewed. Scale review depth to the change's
101
- impact and risk; small, low-risk changes need only a focused review.
102
-
103
- Give the reviewer the original request, constraints, completion criteria, actual
104
- artifacts, and verification evidence. The reviewer must assess both the criteria
105
- and the work, not merely the author's summary. Blocking findings are unmet
106
- required criteria or substantiated, material risks to correctness, security,
107
- data integrity, or usability. Resolve them with fixes or evidence before
108
- proceeding. Separate optional improvements and preferences from blockers.
109
-
110
- Review applies to the reviewed artifact version and context. Reuse it while
111
- both remain applicable; re-review affected areas when changes or new evidence
112
- invalidate it. Review does not replace execution checks. If independent review
113
- is unavailable, stop at the affected gate and report it; self-review does not
114
- satisfy the gate.
115
-
116
- ## 5. Report evidence and stop unproductive loops
117
-
118
- Report what changed, the evidence for completed criteria, and relevant remaining
119
- gaps. Do not present unverified work as complete. Distinguish required checks
120
- from optional broader checks; not running an optional check is not itself a
121
- blocker.
122
-
123
- When retries stop producing new evidence, stop the failing approach and provide
124
- a precise blocker and handoff rather than continuing blindly.
125
-
126
- ## 6. Keep authority explicit
127
-
128
- Work autonomously within the authorized scope. Within existing approvals, carry
129
- the task through implementation, applicable execution checks, and fixes without
130
- pausing for routine confirmation. At a gate, stop only dependent actions and
131
- continue authorized work that does not require crossing it.
132
-
133
- Beyond required reviews, delegate independent tasks when the expected time or
134
- quality benefit outweighs coordination cost. Parallelize implementation only
135
- with non-overlapping ownership and clear interfaces. Keep delegated work within the
136
- same scope and authority; own the integrated result.
137
-
138
- Before destructive or privileged actions, deployment, or shared-state writes,
139
- require explicit approval covering the action and target unless that approval
140
- already exists. A general objective is not approval.
141
-
142
- Ordinary local edits and cleanup of your own disposable artifacts within scope
143
- do not need separate approval. This does not authorize discarding pre-existing
144
- user work or data.
3
+ These are shared decision principles, not a fixed workflow.
4
+ Optimize for a dependable, usable outcome with the least total effort and delay,
5
+ including likely rework and operational consequences. Choose and adapt methods
6
+ to the task, its risks, available capabilities, and evidence.
7
+ Project context refines their application. Agreed outcomes, protection boundaries,
8
+ and explicit authority take precedence over preferred methods.
145
9
 
146
- Preparing a migration, deployment change, or other reviewable artifact does not
147
- authorize applying it to shared systems or persistent application data.
10
+ ## 1. Understand what success means
11
+
12
+ Start from the intended user's goal, the core usage flow, and what makes the
13
+ outcome worth choosing over existing alternatives. Reuse relevant context and
14
+ ground consequential decisions in project evidence or applicable authoritative
15
+ sources. Make assumptions visible when they affect the result.
16
+
17
+ Define success by the outcome users need and the failures they must be protected
18
+ from. Distinguish essential readiness from optional polish, using the intended
19
+ use and agreed constraints rather than a generic notion of completeness.
20
+
21
+ ## 2. Choose and adapt the approach
22
+
23
+ Choose the simplest sufficient solution, considering time to a verified,
24
+ usable result rather than just the first implementation. Use tools, delegation,
25
+ abstractions, and local refactoring where their expected benefit outweighs
26
+ complexity and coordination cost.
27
+
28
+ Resolve routine uncertainty through evidence and reversible progress. Involve
29
+ the user when an unresolved choice has material consequences that cannot be
30
+ safely settled within the agreed intent and authority.
31
+
32
+ Treat methods as replaceable. Improve or replace them when new evidence,
33
+ capabilities, or circumstances offer a better path to the same required outcomes
34
+ and protections.
35
+
36
+ ## 3. Keep changes focused and coherent
37
+
38
+ Let scope follow what the requested outcome genuinely needs, including necessary
39
+ integration and local simplification. Prefer a coherent solution over an
40
+ arbitrarily small diff, while keeping unrelated cleanup and speculative future
41
+ capabilities outside the task.
42
+
43
+ Preserve existing contracts, intentional behavior, and user work unless their
44
+ change is authorized. Remain responsible for the integrated result, including
45
+ work delegated to others.
46
+
47
+ ## 4. Verify outcomes in proportion to consequences
48
+
49
+ Use the affected usage flow and credible, consequential failures to decide what
50
+ needs evidence. Include protections users depend on even when they are not
51
+ visible in the interface.
52
+
53
+ Choose the depth, timing, and combination of testing, review, and release checks
54
+ according to actual impact, uncertainty, recovery cost, and existing evidence.
55
+ Bundle related verification around coherent outcomes and reuse applicable
56
+ evidence across stages, rather than repeating it for each artifact or phase.
57
+
58
+ Give difficult-to-reverse decisions and independently significant risks focused
59
+ attention before the relevant commitment or exposure. Assess reversibility of
60
+ user consequences, not merely the ability to revert code.
61
+
62
+ Use independent review when a fresh perspective materially strengthens assurance,
63
+ or when explicitly required. Scale it to the blind spots and consequential
64
+ mistakes it needs to address.
65
+
66
+ ## 5. Let evidence determine readiness
67
+
68
+ Base conclusions on actual artifacts and observed behavior. Distinguish verified
69
+ results from assumptions, simulated feedback, and work not yet checked. Match
70
+ claims to what the evidence demonstrates.
71
+
72
+ Resolve gaps in required outcomes and substantiated material risks; keep optional
73
+ improvements and preferences separate. Once sufficient evidence supports the
74
+ agreed readiness and applicable requirements, conclude rather than extend the
75
+ work without a likely decision-changing benefit.
76
+
77
+ Use unexpected results to update the diagnosis or approach. When no productive,
78
+ authorized path remains, report the precise blocker and what would resolve it.
79
+ Report what changed, what is supported by evidence, and material remaining gaps.
80
+
81
+ ## 6. Act within clear authority
82
+
83
+ Work autonomously through the authorized outcome, without seeking repeated
84
+ confirmation for routine progress. Keep readiness and authority distinct:
85
+ preparing or validating a change does not authorize applying it to shared
86
+ systems or persistent application data.
87
+
88
+ Reuse approval that covers the action and target. Obtain explicit approval for
89
+ destructive or privileged actions, deployment, and shared-state writes when
90
+ existing authorization does not cover them. A broad objective alone is not
91
+ such approval.
92
+
93
+ Protect user work, data, and secrets, and keep delegated work within the same
94
+ scope and authority. At an unresolved boundary, pause only dependent actions
95
+ and continue useful work that remains authorized.
@@ -14,20 +14,20 @@ Claude Code SSOT(`templates/`, `.claude/`)에서 Codex CLI 프로젝트로 **포
14
14
  templates/codex/
15
15
  ├── README.md # 본 파일
16
16
  ├── AGENTS.md.template # 프로젝트 AGENTS.md (CLAUDE.md에서 생성)
17
- ├── config.toml.template # 프로젝트 .codex/config.toml (hooks + mcp + sandbox)
18
- └── hooks/ # Shell 스크립트 (.claude/hooks/ 재사용 가능)
19
- ├── README.md # 포팅 상태 + Event 매핑
20
- ├── session-start.sh # session_start
21
- └── uncommitted-check.sh # post_tool_use (Bash 한정)
17
+ └── config.toml.template # 프로젝트 .codex/config.toml (hooks + mcp + sandbox)
22
18
  ```
23
19
 
20
+ 훅 스크립트는 이 디렉터리에 없다 — 설치기가 `templates/hooks/session-start.sh` 를 포팅해
21
+ `<project>/.codex/hooks/` 에 놓는다(`src/codex/transform.ts` `HOOK_NAMES`). 자리표시자 사본을 두면
22
+ 설치본에 안 나가는 파일이 리포에 남아 배선을 오독하게 한다(#438).
23
+
24
24
  ## 설치 대상 경로 (`setup-harness.sh --cli=codex` 실행 후)
25
25
 
26
26
  | 템플릿 | 대상 경로 |
27
27
  |--------|-----------|
28
28
  | `AGENTS.md.template` | `<project>/AGENTS.md` |
29
29
  | `config.toml.template` | `<project>/.codex/config.toml` |
30
- | `hooks/*.sh` | `<project>/.codex/hooks/*.sh` |
30
+ | `templates/hooks/session-start.sh` (포팅) | `<project>/.codex/hooks/session-start.sh` |
31
31
  | 사용자 `~/.codex/config.toml`에 trust 등록 추가 | (D4 opt-in 확인) |
32
32
 
33
33
  ## Hook Event 매핑 (Claude → Codex)
@@ -46,7 +46,6 @@ stdin JSON 필드(`session_id`, `cwd`, `tool_name`, `tool_input`, `hook_event_na
46
46
  ## 미구현 (후속 Phase)
47
47
 
48
48
  - `AGENTS.md.template` 본문 — CLAUDE.md에서 변환 (Phase C)
49
- - `hooks/*.sh` 본체 — `.claude/hooks/` 에서 포팅 (Phase C)
50
49
  - `setup-harness.sh --cli=codex` 경로 — Phase D
51
50
  - Trust entry 확인 프롬프트 — Phase D
52
51
  - `--cli=codex` dogfood 검증 — Phase F
@@ -42,7 +42,7 @@ enabled = true
42
42
  # Event 이벤트 매핑: ADR-002 v2 D2 / templates/codex/README.md
43
43
  # ============================================================
44
44
 
45
- # SessionStart: 세션 컨텍스트 + SPEC 재참조 안내 (git pull 없음 — Phase C 포팅 전까지는 placeholder no-op)
45
+ # SessionStart: 세션 컨텍스트 + SPEC 재참조 안내 — 본체는 templates/hooks/session-start.sh 를 설치기가 포팅한다(src/codex/transform.ts HOOK_NAMES)
46
46
  [[hooks.session_start]]
47
47
  name = "session-start"
48
48
  command = ["{PROJECT_DIR}/.codex/hooks/session-start.sh"]
@@ -50,14 +50,6 @@ type = "command"
50
50
  timeout = 10
51
51
  async = false
52
52
 
53
- # PostToolUse: 미커밋 파일 경고 (Bash 한정)
54
- [[hooks.post_tool_use]]
55
- name = "uncommitted-check"
56
- command = ["{PROJECT_DIR}/.codex/hooks/uncommitted-check.sh"]
57
- type = "command"
58
- timeout = 5
59
- async = false
60
-
61
53
  # ============================================================
62
54
  # MCP Servers — .mcp.json에서 포맷 변환
63
55
  # 조건부 포함: Track 따라 setup-harness가 추가
@@ -1,8 +1,11 @@
1
1
  # audit-harness-fit
2
2
 
3
- 현재 합의와 리포의 근거를 기준으로 에이전트 지침·스킬을 정비하고,
4
- `AGENTS.md` / `CLAUDE.md`의 프로젝트 맥락을 채우는 스킬입니다.
5
- 후보 수를 제한하지 않고, 같은 원인을 묶어 영향도순으로 보고합니다.
3
+ 에이전트가 **자율적으로 방법을 선택하면서, 필요한 품질을 지키고, 적정한 시간·비용·
4
+ 사람의 개입으로 사용자 목표를 달성하도록** 지침과 스킬을 정비합니다.
5
+ 현재 합의와 리포의 근거를 기준으로 판단하며, `AGENTS.md` / `CLAUDE.md`의
6
+ 프로젝트 맥락도 채우거나 갱신합니다.
7
+
8
+ > 달성할 결과와 지켜야 할 경계는 분명하게, 작업 방법은 상황에 맞게 선택합니다.
6
9
 
7
10
  ## 디렉터리
8
11
 
@@ -19,33 +22,38 @@ audit-harness-fit/
19
22
  └── scenarios.yaml
20
23
  ```
21
24
 
22
- `SKILL.md`는 진입점입니다. 감사, 검증 설계, 승인된 적용, 맥락 채우기별로
23
- 필요한 참조만 읽도록 구성했습니다. 이 README와 evals는 관리·검토용이며
24
- 일반 실행의 필수 컨텍스트가 아닙니다. 별도 실행 스크립트·모델 API 키·훅은 없습니다.
25
+ `SKILL.md`는 진입점입니다. 감사, 검증 설계, 승인된 적용, 맥락 채우기에 필요한
26
+ 참조만 읽습니다. README와 evals는 관리·검토용이며 일반 실행의 필수 컨텍스트가
27
+ 아닙니다. 별도 실행 스크립트·모델 API 키·훅은 없습니다.
25
28
 
26
29
  ## 하는 일
27
30
 
28
31
  | 영역 | 결과 |
29
32
  |---|---|
30
- | 불필요한 질문·반복 검증 | 실제 지연을 만드는 지시와 조건부 교체 문안, 유효한 근거 재사용 조건 |
31
- | 지침 충돌 | 양쪽 원문, 충돌 상황, 확정된 현재 의도와 실제 구현, 수정안 |
32
- | 과도한 원칙·불필요한 스킬 | 유지·수정·조건 축소·통합·이관·제거·보류 판단과 의존 관계 |
33
- | 긴 결정 사유·히스토리 | 본문·메뉴에는 현재 지시만, 이력은 필요할 때 읽는 별도 문서로 참조 |
34
- | 테스트·상위 모델 위임 | 사용자 사용 장면 → 구현 경로 → 충분한 검증, 필요한 판단만 조건부 위임 |
35
- | 프로젝트 맥락 | 실제 리포로 기존 스캐폴드를 채우거나 갱신하고 미확정 정보는 그대로 표시 |
36
-
37
- 전체 감사는 다섯 감사 영역을 다룹니다. 특정 영역만 요청하면 범위를 확대하지
38
- 않습니다. 문제를 억지로 만들지 않으며, 결과가 많아도 임의의 상위 개수로 자르지
39
- 않습니다. 중요한 미확인 영역과 생략된 상세 내용은 따로 표시합니다.
33
+ | 불필요한 질문·반복 검증 | 실제 마찰을 만드는 지시, 유용한 기본 행동과 추가 확인 조건, 유효한 근거 재사용 |
34
+ | 지침 충돌 | 양쪽 원문과 충돌 상황, 확정된 의도와 실제 구현, 적용 권한에 맞는 수정안 |
35
+ | 지침 가치·자율 선택·모델 적합성 | 유지·수정·조건 축소·보충·통합·이관·제거·보류 판단, 유용한 도구와 맥락 보존 |
36
+ | 긴 결정 사유·히스토리 | 본문과 메뉴에는 현재 지시, 긴 이력은 필요할 때 읽는 별도 문서 |
37
+ | 사용자 결과·검증·실행 경로 | 사용 장면과 핵심 계약별 충분한 증거, 상황에 맞는 검증 시점과 도구·모델·리뷰어 선택 |
38
+ | 프로젝트 맥락 | 실제 리포에 근거한 기존 스캐폴드 보완, 완료 기준과 검증 수단 연결, 미확정 정보 표시 |
39
+
40
+ 전체 감사는 앞의 다섯 영역을 다룹니다. 특정 영역만 요청하면 그 범위를 유지합니다.
41
+ 후보 수는 제한하지 않고 같은 원인을 묶습니다. 모든 중요한 발견을 보존하되,
42
+ 설명의 깊이는 영향과 불확실성에 맞춥니다. 미확인 영역과 생략한 상세 내용은
43
+ 구분해서 표시합니다.
44
+
45
+ 정리만이 답은 아닙니다. 질문의 원인이 실행 위치·명령의 전제·완료 기준의 공백이면
46
+ 근거 있는 맥락을 짧게 보충합니다. 지침의 길이나 검사 횟수보다 실제 결과와
47
+ 총 작업 부담을 기준으로 개선안을 판단합니다.
40
48
 
41
49
  ## 사용
42
50
 
43
- 감사만 수행:
51
+ 감사와 개선안만 요청:
44
52
 
45
53
  ```text
46
- audit-harness-fit으로 현재 적용되는 지침과 Skill을 전체 점검해줘.
47
- 후보 수를 제한하지 말고 중복을 묶어 영향도순으로 정리해줘.
48
- 원문·문제 상황·수정안과 확인/미확인 경로를 보여줘. 파일은 수정하지 마.
54
+ audit-harness-fit으로 현재 적용되는 지침과 스킬을 전체 점검해줘.
55
+ 자율 선택·생산성·품질을 함께 평가하고, 제거할 절차와 보충할 맥락을 판단해줘.
56
+ 후보 수를 제한하지 말고 원문·문제 상황·수정안·확인 범위를 보여줘. 파일은 수정하지 마.
49
57
  ```
50
58
 
51
59
  확인한 제안 적용:
@@ -63,17 +71,38 @@ audit-harness-fit으로 AGENTS.md와 CLAUDE.md의 미완성 프로젝트 맥락
63
71
  핵심 사용 장면과 검증 기준을 연결해줘. 확인되지 않은 내용은 미확정으로 남겨줘.
64
72
  ```
65
73
 
66
- 기존 맥락을 갱신할 때는 "미완성 내용을 채워줘" 대신 "프로젝트 맥락을 최신
67
- 리포와 확정된 결정에 맞춰 갱신해줘"로 요청합니다. 일반 개발 요청마다 전체
68
- 감사를 실행하거나 스캐폴드를 자동으로 다시 채우는 방식은 사용하지 않습니다.
74
+ 기존 맥락 갱신은 "프로젝트 맥락을 최신 리포와 확정된 결정에 맞춰 갱신해줘"로
75
+ 요청합니다. 일반 개발 요청은 기존 개발 흐름을 유지하며, 이 스킬의 설치 자체가
76
+ 전체 감사나 스캐폴드 갱신의 실행 조건은 아닙니다.
77
+
78
+ ## 자율 선택과 모델 발전
79
+
80
+ 검증 시점·묶음 크기·작업 분할·실행 경로는 불확실성, 영향, 피드백 속도와 비용에
81
+ 맞춰 선택합니다. 빠르고 충분한 전체 테스트도, 계약 중심의 선별 검사도 적절한
82
+ 선택이 될 수 있습니다. 여러 핵심 계약을 한 실행으로 입증할 수 있다면 그 증거를
83
+ 재사용하며, 필수 독립 리뷰는 별도로 지킵니다.
84
+
85
+ 모델이 바뀌어도 권한·보안·데이터 보호·정직한 결과 보고는 유지합니다. 프로젝트
86
+ 지식은 사실이 바뀌었는지 확인하고, 특정 모델의 약점을 보완하던 절차는 현재
87
+ 모델·도구와 대표 작업의 증거로 다시 판단합니다. 새 모델이라는 이유만으로
88
+ 절차를 없애거나 전체 감사를 자동 실행하지 않습니다.
89
+
90
+ 효과가 불확실하면 **제안 → 승인된 제한적 시험 → 채택·수정·복원**을 구분합니다.
91
+ 시험은 허용된 격리 환경·도구·데이터·작업 범위에서만 진행합니다. 문서 수정 승인과
92
+ 제품 과제 실행 승인은 다릅니다. 채택 조건까지 이미 승인됐다면 같은 승인을
93
+ 반복해서 요청하지 않습니다. 명백한 중복 정리에는 별도 실험이 필요하지 않습니다.
94
+
95
+ 상위 모델만이 대안은 아닙니다. 기존 경로를 기본값으로 재사용하되, 작업에 맞는
96
+ 도구·전문 에이전트·충분한 저비용 모델·더 유능한 모델·허용된 사람 리뷰어를
97
+ 선택할 수 있습니다. 위임의 이득에는 전달·검토 비용도 포함합니다.
69
98
 
70
99
  ## 배치와 기존 저장소 반영
71
100
 
72
- ZIP 최상위는 `audit-harness-fit/` 하나입니다. 전체 폴더를 사용하는 클라이언트가
73
- 인식하는 스킬 디렉터리에 배치합니다. 클라이언트별 로딩 경로는 실제 설정에서
74
- 확인하고, ZIP을 풀었다는 것만으로 등록·로딩이 완료됐다고 판단하지 않습니다.
101
+ ZIP 최상위는 `audit-harness-fit/` 하나입니다. 사용하는 클라이언트가 인식하는
102
+ 스킬 디렉터리에 전체 폴더를 배치합니다. 로딩 경로는 실제 설정에서 확인하며,
103
+ 압축 해제와 실제 로딩 확인은 구분합니다.
75
104
 
76
- 사용자가 지정한 하네스 저장소의 배포 위치는 다음 형태로 반영할 수 있습니다.
105
+ 하네스 저장소에서는 기존 위치를 유지합니다.
77
106
 
78
107
  ```text
79
108
  templates/skills/audit-harness-fit/SKILL.md
@@ -81,29 +110,33 @@ templates/skills/audit-harness-fit/references/...
81
110
  ```
82
111
 
83
112
  기존 폴더와 먼저 비교하고 로컬 수정사항을 보존합니다. 덮어쓰기만으로 옛 파일이
84
- 삭제되는 것은 아닙니다. 예전 `references/official-criteria.md` 등의 기존 문서는
85
- 이 실행 경로에서 요구하지 않지만, 참조·테스트·이력 보존 필요를 확인하기 전에는
86
- 일괄 삭제하지 않습니다. 배포/로컬 사본이 있다면 저장소의 동기화 방식도 확인합니다.
113
+ 삭제되는 것은 아닙니다. 예전 `references/official-criteria.md` 등 다른 문서는
114
+ 참조·테스트·이력 보존 필요를 확인한 뒤 별도 판단합니다. 배포/로컬 사본이 있다면
115
+ 저장소의 동기화 방식도 확인합니다.
87
116
 
88
- 기본 추천 설치는 기존 자산 ID로 유지하는 통합 작업입니다. 이 ZIP에는 설치기
89
- 코드 변경이 포함되지 않습니다. 설치 후 안내에는 스킬이 실제 설치된 경우에만
90
- "이 스킬로 프로젝트 맥락을 채워 달라"는 문장을 연결하고, 기존 FILL 지시와의
91
- 충돌은 함께 정리합니다. 모든 요청에 붙는 자동 감사 훅은 추가하지 않습니다.
117
+ 기본 추천 설치는 기존 자산 ID를 사용하는 별도 통합 작업입니다. 이 패키지에는
118
+ 설치기 코드 변경이 없습니다. 설치 후 맥락 채우기 안내는 실제 설치 여부에 맞춰
119
+ 연결하고, 기존 FILL 지시와의 충돌은 통합 범위에서 정리합니다. 모든 개발 요청에
120
+ 붙는 자동 감사 훅은 포함하지 않습니다.
92
121
 
93
122
  ## 검토와 제한
94
123
 
95
- [행동 시나리오](evals/scenarios.yaml)는 이 스킬을 개정하거나 특정 동작을 점검할
96
- 때 선택해 쓰는 합성 사례입니다. 별도의 자동 실행기나 실행 결과가 아닙니다.
97
- 일반 프로젝트 작업의 추가 테스트 게이트로 읽거나 실행하지 않습니다.
124
+ [행동 시나리오](evals/scenarios.yaml)는 개정한 동작을 점검할 때 선택하는 합성
125
+ 사례입니다. 자동 실행기나 실행 결과가 아니며, 일반 개발 작업의 추가 게이트로
126
+ 사용하지 않습니다. 권한 위반 방지뿐 아니라 유효한 다른 방법의 선택, 필요한
127
+ 맥락 보충, 충분한 검증의 재사용, 제한적 시험과 품질 보존을 함께 평가합니다.
128
+
129
+ 실제 비교에서는 같은 사용자 결과·완료 기준·필수 보호장치를 유지하고, 완료와
130
+ 품질을 먼저 확인한 뒤 시간·비용·불필요한 질문·재작업을 비교합니다. 지침 분량이나
131
+ 검사 횟수가 줄어도 필요한 증거가 사라졌다면 개선으로 인정하지 않습니다.
98
132
 
99
133
  필수 테스트·독립 리뷰·보안·데이터 보호·배포 승인은 유지합니다. 감사는 읽기
100
- 전용이고, 명시된 범위의 수정만 적용합니다. 상위 모델 사용은 실제 도구·권한·
101
- 데이터·비용 조건이 허용할 때만 가능합니다. 모델의 판단은 테스트 실행 증거가
102
- 아니며, 사용 가능한 리뷰어가 없으면 필수 리뷰를 통과했다고 표시하지 않습니다.
134
+ 전용이고, 적용은 명시된 로컬 범위에 한정합니다. 모델 판단은 테스트 실행 증거가
135
+ 아니며, 사용 가능한 리뷰어가 없으면 필수 리뷰는 미완료로 남깁니다.
103
136
 
104
- 폴더와 문서 형식 확인은 실제 클라이언트 로딩, 에이전트 행동, 설치기 연동 테스트를
105
- 대체하지 않습니다. 이 패키지만으로 GitHub 저장소나 사용자의 설치 환경을 바꾸지
106
- 않습니다.
137
+ 폴더·문서 형식 확인은 실제 클라이언트 로딩, 에이전트 행동, 설치기 연동 검증과
138
+ 다릅니다. 이 패키지를 제공하는 것만으로 GitHub 저장소나 설치 환경이 바뀌지는
139
+ 않습니다. `scenarios.yaml`의 상태는 실제 사례 실행 전까지 `not_executed`입니다.
107
140
 
108
141
  ## 형식 참고
109
142