@uzysjung/agent-harness 26.149.0 → 26.151.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.ko.md +1 -1
- package/README.md +1 -1
- package/dist/{chunk-YSW3OLH4.js → chunk-3QBHZUVB.js} +164 -66
- package/dist/chunk-3QBHZUVB.js.map +1 -0
- package/dist/index.js +397 -293
- package/dist/index.js.map +1 -1
- package/dist/trust-tier-drift.js +5 -1
- package/dist/trust-tier-drift.js.map +1 -1
- package/package.json +1 -1
- package/templates/CLAUDE.md +145 -164
- package/templates/agents/build-error-resolver.md +1 -1
- package/templates/agents/plan-checker.md +1 -1
- package/templates/agents/reviewer.md +4 -5
- package/templates/antigravity/AGENTS.md.template +3 -23
- package/templates/codex/AGENTS.md.template +5 -56
- package/templates/hooks/protect-files.sh +4 -0
- package/templates/hooks/session-start.sh +57 -3
- package/templates/opencode/AGENTS.md.template +4 -52
- package/templates/opencode/opencode.json.template +0 -8
- package/templates/rules/change-management.md +0 -1
- package/templates/rules/cli-development.md +1 -1
- package/templates/rules/doc-governance.md +2 -0
- package/templates/rules/git-policy.md +1 -1
- package/templates/rules/ship-checklist.md +3 -3
- package/templates/rules/test-policy.md +3 -8
- package/templates/settings.json +1 -16
- package/templates/skills/agent-introspection-debugging/SKILL.md +1 -1
- package/templates/skills/audit-harness-fit/README.md +113 -0
- package/templates/skills/audit-harness-fit/SKILL.md +64 -433
- package/templates/skills/audit-harness-fit/evals/scenarios.yaml +222 -0
- package/templates/skills/audit-harness-fit/references/apply.md +66 -0
- package/templates/skills/audit-harness-fit/references/audit.md +160 -0
- package/templates/skills/audit-harness-fit/references/populate.md +74 -0
- package/templates/skills/audit-harness-fit/references/verification.md +123 -0
- package/templates/skills/audit-service-gaps/SKILL.md +6 -7
- package/templates/skills/clear-korean-communication/SKILL.md +8 -13
- package/templates/skills/compaction-handoff/SKILL.md +29 -12
- package/templates/skills/external-model-consult/SKILL.md +13 -24
- package/templates/skills/model-orchestration/SKILL.md +18 -15
- package/templates/skills/natural-korean/SKILL.md +45 -0
- package/templates/skills/north-star/SKILL.md +4 -6
- package/templates/skills/north-star/references/roadmap-method.md +2 -2
- package/templates/skills/{task-brief → objective-brief}/SKILL.md +17 -18
- package/templates/skills/recurrence-prevention/SKILL.md +16 -16
- package/dist/chunk-YSW3OLH4.js.map +0 -1
- package/templates/agents/code-reviewer.md +0 -237
- package/templates/agents/security-reviewer.md +0 -108
- package/templates/hooks/task-brief-nudge.sh +0 -57
- package/templates/skills/audit-harness-fit/references/official-criteria.md +0 -367
- package/templates/skills/continuous-learning-v2/SKILL.md +0 -361
- package/templates/skills/continuous-learning-v2/agents/observer-loop.sh +0 -362
- package/templates/skills/continuous-learning-v2/agents/observer.md +0 -189
- package/templates/skills/continuous-learning-v2/agents/session-guardian.sh +0 -150
- package/templates/skills/continuous-learning-v2/agents/start-observer.sh +0 -252
- package/templates/skills/continuous-learning-v2/config.json +0 -8
- package/templates/skills/continuous-learning-v2/hooks/observe.sh +0 -585
- package/templates/skills/continuous-learning-v2/scripts/detect-project.sh +0 -322
- package/templates/skills/continuous-learning-v2/scripts/instinct-cli.py +0 -1956
- package/templates/skills/continuous-learning-v2/scripts/lib/homunculus-dir.sh +0 -31
- package/templates/skills/continuous-learning-v2/scripts/migrate-homunculus.sh +0 -68
- package/templates/skills/continuous-learning-v2/scripts/test_parse_instinct.py +0 -1420
- package/templates/skills/humanize-korean/SKILL.md +0 -228
- package/templates/skills/spec-scaling/SKILL.md +0 -89
- package/templates/skills/strategic-compact/SKILL.md +0 -145
- package/templates/skills/strategic-compact/suggest-compact.sh +0 -54
|
@@ -1,442 +1,73 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: audit-harness-fit
|
|
3
3
|
description: >-
|
|
4
|
-
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
and waits for approval. Use when the user says any of "하네스 정리", "룰·훅이 밥값을 하는지",
|
|
11
|
-
"CLAUDE.md 다이어트", "상주 컨텍스트 정리", "룰이 너무 많아", "훅이 실제로 뭘 막고 있는지" — or
|
|
12
|
-
"harness audit", "trim my CLAUDE.md", "are my rules earning their keep", "prune the steering
|
|
13
|
-
layer", "why does Claude ignore my rules". Do NOT use it after a specific defect recurred (that is
|
|
14
|
-
recurrence-prevention), or for drift between the product's docs and its code (that is
|
|
15
|
-
audit-service-gaps DRIFT).
|
|
4
|
+
Audit or clean up agent instructions and skills: remove needless questions
|
|
5
|
+
and rechecks, reconcile changed decisions, retire low-value guidance, move
|
|
6
|
+
history to references, and right-size user-journey verification. Also fill
|
|
7
|
+
or refresh AGENTS.md / CLAUDE.md project context from repository evidence.
|
|
8
|
+
Use for harness cleanup, rule conflicts, or context scaffolding, not
|
|
9
|
+
ordinary feature implementation.
|
|
16
10
|
---
|
|
17
11
|
|
|
18
|
-
# Audit Harness Fit
|
|
12
|
+
# Audit Harness Fit
|
|
19
13
|
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
time"*, *"a code review catches something Claude should have known"* — and the natural response is
|
|
23
|
-
another line in CLAUDE.md, another rule file, another hook. Nothing in that loop ever fires in
|
|
24
|
-
reverse. So the layer grows monotonically until it hits the failure the same guidance names:
|
|
25
|
-
*"If your CLAUDE.md is too long, Claude ignores half of it because important rules get lost in the
|
|
26
|
-
noise."*
|
|
14
|
+
Keep agent guidance aligned with confirmed intent, the actual repository,
|
|
15
|
+
and useful current capabilities. Optimize delivery without weakening safeguards.
|
|
27
16
|
|
|
28
|
-
|
|
29
|
-
on each part with **three kinds of evidence only**:
|
|
17
|
+
## Route the request
|
|
30
18
|
|
|
31
|
-
|
|
32
|
-
pruning question. Quotes and sources:
|
|
33
|
-
[references/official-criteria.md](references/official-criteria.md).
|
|
34
|
-
2. **Block and correction logs** — what the enforcement layer actually stopped, and what the
|
|
35
|
-
human actually had to correct.
|
|
36
|
-
3. **Measurement** — item counts and byte sizes per surface, taken the same way twice so before
|
|
37
|
-
and after are comparable.
|
|
38
|
-
|
|
39
|
-
The three are collection channels, not equals. **The published checklist ranks first**: it is the
|
|
40
|
-
only one that does not depend on this project's local sample, so a section that the checklist
|
|
41
|
-
already answers is not re-argued from logs or counts. Logs and measurement decide what the
|
|
42
|
-
checklist leaves open.
|
|
43
|
-
|
|
44
|
-
**Opinion is not one of the three.** "This rule feels important" and "this rule feels like bloat"
|
|
45
|
-
are the same evidence class, and a verdict that pits one against the other is a coin flip wearing
|
|
46
|
-
a report's clothes. If none of the three applies to a section, the verdict is `unjudged`, and it
|
|
47
|
-
stays exactly as it is.
|
|
48
|
-
|
|
49
|
-
## When to use
|
|
50
|
-
|
|
51
|
-
- The resident layer has grown across many sessions and nobody has ever removed anything.
|
|
52
|
-
- Claude keeps violating a rule that is plainly written down — the diagnostic the docs give for
|
|
53
|
-
that symptom is *file length*, not rule wording.
|
|
54
|
-
- You are about to add another rule and want to know what the existing ones are doing first.
|
|
55
|
-
- A model upgrade landed and some instructions may now be scaffolding for a weakness that is gone.
|
|
56
|
-
|
|
57
|
-
Not for: a defect that just recurred — `recurrence-prevention` owns that, and it moves *one*
|
|
58
|
-
countermeasure up a ladder rather than re-judging the whole layer. Not for mismatch between the
|
|
59
|
-
product's documentation and the product's code — that is `audit-service-gaps` in DRIFT mode.
|
|
60
|
-
|
|
61
|
-
---
|
|
62
|
-
|
|
63
|
-
## Stage 1 — INVENTORY (what actually loads every session)
|
|
64
|
-
|
|
65
|
-
Enumerate the resident surfaces before judging any of them. Resident means loaded at session
|
|
66
|
-
start whether or not it gets used:
|
|
67
|
-
|
|
68
|
-
| Surface | Where it lives | Resident? |
|
|
69
|
-
|---|---|---|
|
|
70
|
-
| Project anchor | `CLAUDE.md` / `AGENTS.md` at the repo root (and parent directories) | Full text, every session |
|
|
71
|
-
| Project anchor 2 | `.claude/CLAUDE.md` — a second file, not an alias of the first | Full text, every session |
|
|
72
|
-
| User anchor | the same filenames under the home config dir (`$HOME/.claude/`) | Full text — a separate scope, not a parent directory |
|
|
73
|
-
| Imports | every `@path` reachable from any of those anchors, up to four hops | Full text — imports organize, they do not reduce |
|
|
74
|
-
| Auto-memory | `MEMORY.md`, written by the agent to itself | Full text; often the single largest item |
|
|
75
|
-
| Rules | `.claude/rules/*.md` | Full text if no `paths:` frontmatter; on match if scoped |
|
|
76
|
-
| Skills | `.claude/skills/*/SKILL.md` | **Name + description only**; the body loads on trigger |
|
|
77
|
-
| Subagents | `.claude/agents/*.md` | Description preloaded, same as skills |
|
|
78
|
-
| Hooks | `settings.json` hook entries + their scripts | **Zero**, unless the hook writes to stdout |
|
|
79
|
-
| Permissions | `permissions.allow` / `ask` / `deny` | Enforcement, not context |
|
|
80
|
-
|
|
81
|
-
Then measure. Use whatever the project already provides; if it provides nothing, plain shell is
|
|
82
|
-
enough and portable. Measure every row in the same unit — **bytes** — or the rows cannot be added
|
|
83
|
-
up, and "half the layer is rules" becomes a guess:
|
|
84
|
-
|
|
85
|
-
```bash
|
|
86
|
-
have() { for f in "$@"; do [ -f "$f" ] && printf '%s\n' "$f"; done; }
|
|
87
|
-
|
|
88
|
-
# ⓐ every anchor that exists, plus one hop of @imports resolved next to the file that declared them
|
|
89
|
-
anchors=$(have CLAUDE.md .claude/CLAUDE.md AGENTS.md "$HOME/.claude/CLAUDE.md")
|
|
90
|
-
imports=$(for a in $anchors; do
|
|
91
|
-
grep -o '@[^[:space:])]*' "$a" |
|
|
92
|
-
sed -e "s|^@~|$HOME|" -e "s|^@/|/|" -e "s|^@|$(dirname "$a")/|"
|
|
93
|
-
done)
|
|
94
|
-
have $anchors $imports | xargs wc -c # ÷ 4 ≈ tokens
|
|
95
|
-
|
|
96
|
-
# ⓑ rules — count them, then size only the ones without `paths:` frontmatter
|
|
97
|
-
find .claude/rules -name '*.md' | wc -l
|
|
98
|
-
grep -L '^paths:' .claude/rules/*.md | xargs wc -c
|
|
99
|
-
|
|
100
|
-
# ⓒ auto-memory — one home dir holds every project's, so keep the ones naming this project
|
|
101
|
-
find . "$HOME/.claude" -name 'MEMORY.md' 2>/dev/null |
|
|
102
|
-
grep -e '^\./' -e "$(basename "$PWD")" | xargs wc -c
|
|
103
|
-
|
|
104
|
-
# ⓓ for skills and subagents the descriptor is the resident part — size the frontmatter, not the file
|
|
105
|
-
awk 'FNR==1{n=0} /^---$/{n++;next} n==1' .claude/skills/*/SKILL.md | wc -c
|
|
106
|
-
awk 'FNR==1{n=0} /^---$/{n++;next} n==1' .claude/agents/*.md | wc -c
|
|
107
|
-
```
|
|
108
|
-
|
|
109
|
-
Three traps that make an inventory wrong rather than incomplete:
|
|
110
|
-
|
|
111
|
-
- **Two copies of the same name.** A repo that *ships* a harness has a development copy and a
|
|
112
|
-
distributed copy of the same filenames. Measure the one the session actually loads, and say
|
|
113
|
-
which one you measured.
|
|
114
|
-
- **`paths:` frontmatter changes the answer.** A rule with a `paths:` list is not resident; a rule
|
|
115
|
-
without one is. Read the frontmatter, do not assume.
|
|
116
|
-
- **Imports are not free.** Splitting a long anchor into `@imports` improves organization and
|
|
117
|
-
changes the resident total by nothing.
|
|
118
|
-
|
|
119
|
-
Record the numbers as bytes per surface plus one total. Every later claim about "smaller" has to
|
|
120
|
-
point back at them, and a total that quietly drops a surface makes every percentage after it wrong.
|
|
121
|
-
|
|
122
|
-
## Stage 2 — EVIDENCE (what the layer actually did)
|
|
123
|
-
|
|
124
|
-
For each resident item, look for a trace that it did work:
|
|
125
|
-
|
|
126
|
-
- **Block logs.** Blocking hooks that append one line per block (`.uzys-agent-harness/hook-blocks.log`
|
|
127
|
-
where this harness is installed) give the only direct data on what enforcement actually caught.
|
|
128
|
-
- **Correction history.** `git log` on the steering files themselves, plus the commits that
|
|
129
|
-
*followed* a rule's introduction: was the mistake it targets absent afterward, or does it recur?
|
|
130
|
-
- **The maintainer's own record.** Issue threads, postmortems, memory files — a rule created after
|
|
131
|
-
a real incident has a citation; a rule created out of caution does not.
|
|
132
|
-
|
|
133
|
-
**A log with zero lines is not an acquittal.** It has two readings that data alone cannot separate:
|
|
134
|
-
nothing needed blocking, or the log was born last week / the hook never fired / the hook never wired
|
|
135
|
-
up. Report `no sample` and go find a second signal — when the hook was added, whether its matcher can
|
|
136
|
-
ever match, whether the file exists at all. A hook whose matcher cannot match anything is not
|
|
137
|
-
"quietly effective", it is dead wiring, and that is a finding in its own right.
|
|
138
|
-
|
|
139
|
-
The same asymmetry runs the other way: a log line proves the hook fired, not that the block was
|
|
140
|
-
*correct*. Read the blocked targets. A block on a path the maintainer intended to edit is a false
|
|
141
|
-
positive, and false positives are the cost side of the enforcement ledger.
|
|
142
|
-
|
|
143
|
-
## Stage 3 — VERDICT (rule each section against the published checklist)
|
|
144
|
-
|
|
145
|
-
The unit is the **section**, not the file. Files are usually mixed — one paragraph carrying a real
|
|
146
|
-
project fact, three carrying things any competent model already does.
|
|
147
|
-
|
|
148
|
-
Work through four steps in order and stop at the first one that answers. Steps 1 and 2 are
|
|
149
|
-
checklist lookups rather than judgment calls, and they settle most sections:
|
|
150
|
-
|
|
151
|
-
1. **Include categories** — does the section map to one of the five documented categories?
|
|
152
|
-
2. **Exclude list** — is it one of the four things documented as not worth including?
|
|
153
|
-
3. **Form** — right content, wrong wording: specificity and consistency.
|
|
154
|
-
4. **Pruning question** — for whatever steps 1–3 leave open.
|
|
155
|
-
|
|
156
|
-
Record the category (or the exclusion) beside each section, so a reader can re-derive the verdict
|
|
157
|
-
without re-reading the section.
|
|
158
|
-
|
|
159
|
-
### Step 1 — map every section to an include category
|
|
160
|
-
|
|
161
|
-
| Category | Published wording | What lands here |
|
|
162
|
-
|---|---|---|
|
|
163
|
-
| Commands | *"Commands — how to build, test, lint, and run locally"* | invocations the model cannot guess from the repo |
|
|
164
|
-
| Conventions | *"Conventions — naming, error handling, file layout, and 'we use X, not Y'"* | choices that differ from the language default |
|
|
165
|
-
| Architecture | *"Architecture in three sentences — what the major pieces are"* | the shape of the system, not a tour of it |
|
|
166
|
-
| Hard constraints | *"Hard constraints — for example, 'never write to the production database'"* | what must never happen |
|
|
167
|
-
| Known gotchas | *"Known gotchas — the issues every new engineer trips on"* | non-obvious behavior that already cost someone a day |
|
|
168
|
-
|
|
169
|
-
A section mapping to none of the five goes to step 2, not straight to `delete`. A section mapping to
|
|
170
|
-
two is usually two sections. Note the tension the same source states about the third row: the
|
|
171
|
-
vendor's own trim heuristic *"cuts content Claude can derive from the codebase, such as directory
|
|
172
|
-
layouts, dependency lists, and architecture overviews"* — three sentences of shape belong resident,
|
|
173
|
-
an architecture overview does not.
|
|
174
|
-
|
|
175
|
-
The include/exclude table states the same split from the other side, and its exclude column does
|
|
176
|
-
most of the cutting: *"Anything Claude can figure out by reading code"*, *"Standard language
|
|
177
|
-
conventions Claude already knows"*, *"Detailed API documentation (link to docs instead)"*,
|
|
178
|
-
*"Information that changes frequently"*, *"Long explanations or tutorials"*, *"File-by-file
|
|
179
|
-
descriptions of the codebase"*, *"Self-evident practices like "write clean code""*.
|
|
180
|
-
|
|
181
|
-
### Step 2 — check the four named exclusions
|
|
182
|
-
|
|
183
|
-
Four things are named as not worth including. Each has a mechanical check, so this step produces a
|
|
184
|
-
count rather than an opinion:
|
|
185
|
-
|
|
186
|
-
| Exclusion (published wording) | Where it hides |
|
|
187
|
-
|---|---|
|
|
188
|
-
| *"Changelogs or history"* | dated lines, version tags, "as of", "used to", migration notes |
|
|
189
|
-
| *"Full API documentation (Claude can read the code directly)"* + *"Anything that is already obvious from the file tree"* | directory trees, file-by-file lists, exported-symbol lists |
|
|
190
|
-
| *"Information that changes frequently"* | counts, versions, "currently N of M" — anything one release invalidates |
|
|
191
|
-
| *"Aspirational rules the team does not actually follow"* | check the repository's own history: does it obey the rule? |
|
|
192
|
-
|
|
193
|
-
A hit on this list is a `delete` or a `relocate`, never a `keep`. Derivable content in particular is
|
|
194
|
-
derivable *by the model, on demand* — it does not need to be resident.
|
|
195
|
-
|
|
196
|
-
### Step 3 — form: specificity and consistency
|
|
197
|
-
|
|
198
|
-
Right content in the wrong form is a `rewrite`, not a `delete`:
|
|
199
|
-
|
|
200
|
-
> "**Specificity**: write instructions that are concrete enough to verify. For example:
|
|
201
|
-
>
|
|
202
|
-
> * "Use 2-space indentation" instead of "Format code properly"
|
|
203
|
-
> * "Run `npm test` before committing" instead of "Test your changes"
|
|
204
|
-
> * "API handlers live in `src/api/handlers/`" instead of "Keep files organized""
|
|
205
|
-
|
|
206
|
-
> "**Consistency**: if two rules contradict each other, Claude may pick one arbitrarily. Review
|
|
207
|
-
> your CLAUDE.md files, nested CLAUDE.md files in subdirectories, and [`.claude/rules/`]
|
|
208
|
-
> periodically to remove outdated or conflicting instructions."
|
|
209
|
-
|
|
210
|
-
Conflicts deserve their own sweep across every anchor and rule file at once: two sections that
|
|
211
|
-
contradict each other are worse than either alone, because the model may follow either one on any
|
|
212
|
-
given session.
|
|
213
|
-
|
|
214
|
-
### Step 4 — size, and the pruning question
|
|
215
|
-
|
|
216
|
-
For whatever steps 1–3 leave open, the published question decides:
|
|
217
|
-
|
|
218
|
-
> "Keep it concise. For each line, ask: *"Would removing this cause Claude to make mistakes?"* If
|
|
219
|
-
> not, cut it. Bloated CLAUDE.md files cause Claude to ignore your actual instructions!"
|
|
220
|
-
|
|
221
|
-
Anything the model does correctly without the instruction is a no-op that still costs adherence
|
|
222
|
-
from the rules around it — that is a `delete` even when nothing else flagged it.
|
|
223
|
-
|
|
224
|
-
The published size figure is **200 lines per CLAUDE.md file**: *"Longer files consume more context
|
|
225
|
-
and reduce adherence."* Two things it does not mean:
|
|
226
|
-
|
|
227
|
-
- It is not a hard cut-off — *"CLAUDE.md files are loaded in full regardless of length, though
|
|
228
|
-
shorter files produce better adherence."* Report the overage as a number, not as a failure.
|
|
229
|
-
- There is **no published budget for the number of rule files, and none for hooks.** Reporting
|
|
230
|
-
"too many rules" as a criterion is inventing one. Rule each file on the same checklist and report
|
|
231
|
-
the total in bytes.
|
|
232
|
-
|
|
233
|
-
### The generation lint
|
|
234
|
-
|
|
235
|
-
Recent guidance names prompt patterns that were useful for older models and now actively cost
|
|
236
|
-
tokens or quality — they survive in steering layers as legacy scaffolding, so look for them by name:
|
|
237
|
-
|
|
238
|
-
| Pattern to flag | Why it is now a cost |
|
|
239
|
-
|---|---|
|
|
240
|
-
| Explicit verification instructions ("add a final verification step", "use a subagent to verify") | The model verifies its own work unprompted; the instruction causes over-verification |
|
|
241
|
-
| Re-check instructions ("double-check your answer", "re-verify before responding") | Compounds with behavior the model already has — cost without quality |
|
|
242
|
-
| Severity suppression in review prompts ("only report high-severity issues", "be conservative") | Followed literally: the review reports less. Ask for everything, filter in a separate pass |
|
|
243
|
-
| Rules telling the model not to think or not to reason, especially naming thinking tags | Increases tag leakage — the documented effect is the opposite of the intent |
|
|
244
|
-
| Long stacks of prohibitions | Positive examples of the wanted style outperform instructions about what not to do |
|
|
245
|
-
| Aspirational rules nobody follows | Documented as "not worth including"; also teaches that rules are optional |
|
|
246
|
-
|
|
247
|
-
A flag is a candidate, not a verdict. Confirm it against the section's evidence before ruling.
|
|
248
|
-
|
|
249
|
-
### When the model underneath changes — ablate, then re-earn
|
|
250
|
-
|
|
251
|
-
A steering layer accumulates corrections aimed at whichever model was current when each line was
|
|
252
|
-
written. Those lines do not expire on their own. After the project moves to a newer model they keep
|
|
253
|
-
charging adherence to correct mistakes it no longer makes, and the layer reads as a record of past
|
|
254
|
-
model weaknesses rather than of this project.
|
|
255
|
-
|
|
256
|
-
The reset is deliberate rather than gradual: **take the accumulated instructions out, do the work,
|
|
257
|
-
and add back only what an observed, repeated mistake demands.** A line earns its place by a failure
|
|
258
|
-
someone watched happen on the model in use — never by having been true of an older one. This is the
|
|
259
|
-
same bar the vendor sets for writing a steering line at all, applied at the moment the model
|
|
260
|
-
underneath changes.
|
|
261
|
-
|
|
262
|
-
The same reasoning bounds how much method to specify. Instructions that dictate *how* a capable
|
|
263
|
-
model reaches a result cap the result at the author's plan, because the model has to follow them
|
|
264
|
-
literally even when it sees further. State the goal, the constraints that genuinely must hold, and
|
|
265
|
-
how the work will be judged — then leave the method open. Pin down a specific method only where one
|
|
266
|
-
is actually required: an external contract, a boundary that must not be crossed, or a tool the model
|
|
267
|
-
cannot discover on its own (a script this harness installed, for instance).
|
|
268
|
-
|
|
269
|
-
| Pattern | Why it costs |
|
|
270
|
-
|---|---|
|
|
271
|
-
| Instruction carried over from an older model, with no observed failure on the current one | Pure adherence tax — it dilutes the lines that do matter |
|
|
272
|
-
| Step-by-step scaffolding for work the model can plan itself | Caps the outcome at the author's plan and hides better approaches |
|
|
273
|
-
| Method pinned down where only the outcome matters | Same cost, and it goes stale when the tooling changes |
|
|
274
|
-
|
|
275
|
-
Both patterns rule `delete` when nothing in Stage 2's evidence names a failure they prevented.
|
|
276
|
-
|
|
277
|
-
### Assign exactly one verdict per section
|
|
278
|
-
|
|
279
|
-
- **keep** — maps to an include category, is off the exclude list, and belongs resident.
|
|
280
|
-
- **rewrite** — right content, wrong form: vague where it should be concrete ("format properly" →
|
|
281
|
-
"use 2-space indentation"), or contradicting another section.
|
|
282
|
-
- **relocate** — right content, wrong layer. Stage 4 decides where.
|
|
283
|
-
- **delete** — on the exclude list, derivable, already-known, generation lint confirmed, or dead
|
|
284
|
-
wiring.
|
|
285
|
-
- **unjudged** — the checklist does not reach it and none of the three evidence kinds applies.
|
|
286
|
-
Leave it alone and say so.
|
|
287
|
-
|
|
288
|
-
## Stage 4 — RELOCATE (right content, wrong layer)
|
|
289
|
-
|
|
290
|
-
Most of what a bloated steering layer holds is not wrong — it is filed in the layer that cannot
|
|
291
|
-
enforce it and charges rent for trying.
|
|
292
|
-
|
|
293
|
-
| What it is | Where it belongs | Why |
|
|
294
|
-
|---|---|---|
|
|
295
|
-
| Multi-step procedure, playbook, checklist | **Skill** | Loads on demand; the descriptor is the only resident cost |
|
|
296
|
-
| Instruction that only matters for part of the tree | **Path-scoped rule** (`paths:` frontmatter) | Loads when matching files are touched, not every session |
|
|
297
|
-
| Must happen every time, no exceptions (format on save, block a path) | **Hook** | Prose is advisory; hooks are deterministic and fire regardless of what the model decides |
|
|
298
|
-
| Hard allow/deny boundary on tools, commands, paths | **Permission rule** | Documented as the enforcement layer for boundaries; a hook filter is best-effort and fails open on unparseable input |
|
|
299
|
-
| A fact the code already states, or should | **Code, test, or generated doc** | Derived facts do not drift; copied facts do |
|
|
300
|
-
| Dynamic per-session context (recent commits, open issues) | **SessionStart hook** | Static context belongs in the anchor; only scripted, changing context justifies a hook |
|
|
301
|
-
| A system Claude keeps re-reading or cannot see at all; a setup a second repo needs too | **MCP server** or **plugin** | Connect or package the capability instead of narrating it in prose that loads every session |
|
|
302
|
-
| A side task whose output would flood the main conversation | **Subagent** | Runs in its own context; only the result comes back |
|
|
303
|
-
| Nothing depends on it | **Delete** | |
|
|
304
|
-
|
|
305
|
-
Two directions that look symmetric and are not: hooks can tighten what permission rules allow but
|
|
306
|
-
never loosen it, and a prompt instruction is not on the enforcement list at all — it shapes what
|
|
307
|
-
the model attempts, so pair it with one of the two real mechanisms rather than shipping it alone.
|
|
308
|
-
|
|
309
|
-
Relocation is not free either. A hook adds a shell dependency and an administrative surface; a
|
|
310
|
-
skill adds a descriptor to every session. Say what the move costs, not only what it saves.
|
|
311
|
-
|
|
312
|
-
### The reverse move — a skill that never fires
|
|
313
|
-
|
|
314
|
-
Moving a procedure into a skill only pays off if the skill actually loads. Skills load when the model
|
|
315
|
-
judges them relevant to the prompt, which is enough for task-shaped skills ("review this UI") and
|
|
316
|
-
not enough for skills meant to apply to *every* answer or *every* delegation. Those need one resident
|
|
317
|
-
line saying when they apply; without it the skill is installed, costs a descriptor every session, and
|
|
318
|
-
never runs.
|
|
319
|
-
|
|
320
|
-
Write the line only for skills this project actually has — a pointer to an uninstalled skill is a
|
|
321
|
-
dead reference, and this audit exists to remove those, not to add them. Check the install first:
|
|
322
|
-
|
|
323
|
-
| Skill, where installed | The resident line it needs |
|
|
324
|
-
|---|---|
|
|
325
|
-
| `clear-korean-communication` | It applies to every answer, report, and approval request — not only at the moment approval is asked for |
|
|
326
|
-
| `task-brief` | Incoming work requests are normalized into the brief shape before work starts, and the filled-in brief is shown to the user |
|
|
327
|
-
| `model-orchestration` | Delegation follows it — which model and which effort each lane gets is its call, not an ad-hoc pick |
|
|
328
|
-
|
|
329
|
-
One line each. The skill body holds the procedure; the resident line carries only *when it applies*,
|
|
330
|
-
which is the part the model cannot infer from a descriptor.
|
|
331
|
-
|
|
332
|
-
## Stage 5 — APPLY (propose; the human decides)
|
|
333
|
-
|
|
334
|
-
Default output is a **proposal**, not an edit. Present it as one table — section, category or
|
|
335
|
-
exclusion, verdict, evidence, destination — with before/after measurements from Stage 1, then stop.
|
|
336
|
-
|
|
337
|
-
Apply only what was approved, and keep the applied change checkable:
|
|
338
|
-
|
|
339
|
-
- One coherent commit, so the removal can be reverted as a unit.
|
|
340
|
-
- **Deletions are reversible in version control and nowhere else.** Before deleting a rule, check
|
|
341
|
-
whether a test, gate, or script reads that file by path — a gate that greps for a removed file
|
|
342
|
-
turns green by finding nothing.
|
|
343
|
-
- After applying, the honest verification is behavioral: the guidance's own instruction is to
|
|
344
|
-
*"test changes by observing whether Claude's behavior actually shifts."* Say plainly that the
|
|
345
|
-
effect is unverified until that observation exists. A smaller token count is not evidence that
|
|
346
|
-
the layer got better.
|
|
347
|
-
|
|
348
|
-
Never widen the audit into a rewrite of the project's conventions. This skill decides what loads,
|
|
349
|
-
not what the team believes.
|
|
350
|
-
|
|
351
|
-
## Success criteria for a finished audit
|
|
352
|
-
|
|
353
|
-
A run is finished when all five hold. Each is settled by a command or by counting rows — "the layer
|
|
354
|
-
reads tighter now" is not on the list:
|
|
355
|
-
|
|
356
|
-
| # | Criterion | How it is checked |
|
|
19
|
+
| Requested work | Read | Behavior |
|
|
357
20
|
|---|---|---|
|
|
358
|
-
|
|
|
359
|
-
|
|
|
360
|
-
|
|
|
361
|
-
|
|
|
362
|
-
|
|
363
|
-
|
|
364
|
-
|
|
365
|
-
|
|
366
|
-
|
|
367
|
-
|
|
368
|
-
|
|
369
|
-
|
|
370
|
-
|
|
371
|
-
|
|
372
|
-
|
|
373
|
-
|
|
374
|
-
|
|
375
|
-
|
|
376
|
-
|
|
377
|
-
|
|
378
|
-
|
|
379
|
-
|
|
380
|
-
|
|
381
|
-
|
|
382
|
-
|
|
383
|
-
|
|
384
|
-
|
|
385
|
-
|
|
386
|
-
|
|
387
|
-
|
|
388
|
-
|
|
389
|
-
|
|
390
|
-
|
|
391
|
-
|
|
392
|
-
|
|
393
|
-
|
|
394
|
-
|
|
395
|
-
|
|
396
|
-
|
|
397
|
-
|
|
398
|
-
|
|
399
|
-
|
|
400
|
-
|
|
401
|
-
|
|
402
|
-
|
|
403
|
-
|
|
404
|
-
|
|
405
|
-
|
|
406
|
-
|
|
407
|
-
|
|
408
|
-
|
|
409
|
-
|
|
410
|
-
|
|
411
|
-
"never edit `.env`" line stays as prose *and* keeps its hook, since prose alone is not a boundary.
|
|
412
|
-
The dead-wired hook is deleted; the never-fired one is left in place with its status recorded as
|
|
413
|
-
unknown, because deleting on absence of evidence is the mistake this stage exists to avoid.
|
|
414
|
-
|
|
415
|
-
**APPLY** — proposal table presented; maintainer approves the deletions, defers the skill
|
|
416
|
-
extraction. Reported as: −565 resident lines, exclude-list greps ⓐ–ⓓ clean, project test and lint
|
|
417
|
-
commands green, 1 dead hook removed, 1 false-positive block identified; **behavioral effect
|
|
418
|
-
unverified** until the next few sessions are observed.
|
|
419
|
-
|
|
420
|
-
## Output, side effects, and stop conditions
|
|
421
|
-
|
|
422
|
-
- **Output** — the inventory with its measurement method, the evidence table (including `no
|
|
423
|
-
sample` entries), one row per section with its category-or-exclusion and verdict, the relocation
|
|
424
|
-
plan with costs, the before/after numbers, and the five success criteria with their check results.
|
|
425
|
-
- **Side effects** — this skill proposes; it edits only what was approved. Never touch
|
|
426
|
-
`permissions` or hook configuration without explicit approval: those change what the agent is
|
|
427
|
-
allowed to do, not merely what it reads.
|
|
428
|
-
- **Stop** when the project's steering layer is spread across copies you cannot tell apart, or
|
|
429
|
-
when the only available judgment is preference. An audit that ranks sections by taste produces a
|
|
430
|
-
confident list of changes that no one can defend later.
|
|
431
|
-
|
|
432
|
-
## Cross-references (don't duplicate)
|
|
433
|
-
|
|
434
|
-
- **`recurrence-prevention`** — opposite direction: it escalates one countermeasure after a
|
|
435
|
-
specific defect returned. This skill audits the layer at rest; it does not decide whether a
|
|
436
|
-
given incident deserves a new rule.
|
|
437
|
-
- **`audit-service-gaps`** — audits the product against its target state; DRIFT mode covers
|
|
438
|
-
doc-vs-code mismatch in the product. This skill audits the agent's own steering layer.
|
|
439
|
-
- **`north-star`** — where a project's stated direction lives; a rule that no longer serves it is
|
|
440
|
-
a candidate for deletion, but the direction itself is set there, not here.
|
|
441
|
-
- **[references/official-criteria.md](references/official-criteria.md)** — the published quotes
|
|
442
|
-
behind every criterion above, with sources. Read it before ruling on a contested section.
|
|
21
|
+
| Audit, reconcile, or propose cleanup | [Audit](references/audit.md) | Read-only findings and proposed edits |
|
|
22
|
+
| A full audit, or testing / usage-scene / model-routing advice | [Verification](references/verification.md), plus Audit for findings | Propose scenario-based implementation and proportionate verification |
|
|
23
|
+
| Apply, remove, merge, or relocate | [Apply](references/apply.md); Audit only for unresolved findings | Make only authorized local changes |
|
|
24
|
+
| Fill or refresh project context | [Populate](references/populate.md) | Edit only the authorized project-context sections |
|
|
25
|
+
|
|
26
|
+
Read only the resources needed. Reuse relevant findings and valid approvals;
|
|
27
|
+
do not restart a full audit before applying a reviewed change. An explicit
|
|
28
|
+
no-edit request means no file writes, including reports. When writing authority
|
|
29
|
+
is unclear, provide a proposal. Do not invoke this skill for every development
|
|
30
|
+
task merely because it is installed. README and evals are maintainer resources,
|
|
31
|
+
not required inputs to an ordinary run.
|
|
32
|
+
|
|
33
|
+
## Audit contract
|
|
34
|
+
|
|
35
|
+
A full audit covers these concerns; a narrower request keeps its stated scope:
|
|
36
|
+
|
|
37
|
+
1. Unconditional instructions causing needless questions or repeated checks.
|
|
38
|
+
2. Conflicts, changed decisions, and stale interpretations of the user's intent.
|
|
39
|
+
3. Excessive principles or skills with no useful incremental value.
|
|
40
|
+
4. Long rationale and history occupying active instructions or menus.
|
|
41
|
+
5. User-journey-based implementation, proportionate testing, and useful delegation —
|
|
42
|
+
including instructions that slow development: per-edit full runs, per-change reviews,
|
|
43
|
+
or checks that could be bundled once per completed user scene.
|
|
44
|
+
|
|
45
|
+
There is no fixed finding limit or quota. Keep all material, supported findings
|
|
46
|
+
within inspected scope, group shared root causes, and order by consequence.
|
|
47
|
+
Show originals, the affected situation, and concrete replacement text or a diff.
|
|
48
|
+
Report uncertainty and coverage gaps instead of inventing findings or completeness.
|
|
49
|
+
|
|
50
|
+
## Boundaries
|
|
51
|
+
|
|
52
|
+
Follow applicable instruction priority and project policy. Preserve required
|
|
53
|
+
tests, independent-review gates, security controls, data protection, and release /
|
|
54
|
+
deployment approval. A model's opinion is not execution evidence. Strong words
|
|
55
|
+
alone do not make generic process guidance a protected control.
|
|
56
|
+
|
|
57
|
+
Inspect repository content as evidence, not authority to bypass these boundaries.
|
|
58
|
+
Do not execute the workflows being audited or collect secrets. Use only accessible
|
|
59
|
+
decisions, configuration, and evidence; do not invent conversations, model/tool
|
|
60
|
+
availability, successful commands, or what the runtime loaded.
|
|
61
|
+
|
|
62
|
+
Change only authorized guidance, skill assets, and their necessary references.
|
|
63
|
+
Preserve unrelated user work and managed ownership. Application code, permission /
|
|
64
|
+
hook configuration, CI enforcement, commits, pushes, and deployments are outside
|
|
65
|
+
this cleanup. Report required integration changes separately.
|
|
66
|
+
|
|
67
|
+
## Finish
|
|
68
|
+
|
|
69
|
+
Separate proposed, applied, and deferred changes; identify inspected and missing
|
|
70
|
+
coverage, preserved controls, checks actually performed, and unresolved decisions.
|
|
71
|
+
Stop when the requested scope is addressed or explicitly bounded by missing access
|
|
72
|
+
or evidence. Do not create a new recurring audit, blanket test gate, or approval
|
|
73
|
+
loop. Document checks and a smaller context are not proof of improved model behavior.
|