@uzysjung/agent-harness 26.156.0 → 26.158.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -17,7 +17,7 @@ import {
17
17
  init_esm_shims,
18
18
  residentCost,
19
19
  resolveBundleRoot
20
- } from "./chunk-PPJUFG53.js";
20
+ } from "./chunk-HHCNPPGS.js";
21
21
 
22
22
  // src/trust-tier-drift.ts
23
23
  init_esm_shims();
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@uzysjung/agent-harness",
3
- "version": "26.156.0",
3
+ "version": "26.158.0",
4
4
  "description": "Curate vetted AI-coding skills & plugins by your tech stack — install only what you need, across Claude Code, Codex, OpenCode & Antigravity",
5
5
  "type": "module",
6
6
  "publishConfig": {
@@ -6,6 +6,16 @@ belongs to the Delivery rule.
6
6
  - An ordinary change gets the repository's baseline CI and regression across the affected scope;
7
7
  independent verification only where the Delivery rule requires it. A high-risk one widens that in
8
8
  proportion to what it touches and what its failure would cost.
9
+ - Preserve required outcomes with the least burdensome reliable protection.
10
+ Prefer fixing the cause and reusing or improving existing mechanisms within scope;
11
+ no new guard is needed when those suffice. Judge safeguards by credible risk,
12
+ the assurance they add, and total effort, not request emphasis, incident counts,
13
+ or the number or subject of checks. Tests of safeguards follow the same standard.
14
+ Reuse valid evidence and recheck affected claims when relevant conditions change.
15
+ Within project policy, adapt methods to demonstrated model and tool capabilities
16
+ while preserving acceptance criteria, required gates, and authority boundaries.
17
+ Finish optional verification when required outcomes have sufficient evidence
18
+ and residual risk is within the project's accepted limits.
9
19
  - High-risk includes at least authentication, authorization, payments and settlement, personal
10
20
  data, data integrity, concurrency, state transitions, and migrations.
11
21
  - Full regression, full E2E, full mutation, and periodic security scanning belong to the CI/CD
@@ -16,5 +26,7 @@ belongs to the Delivery rule.
16
26
  the result. Otherwise, use explicit test doubles or contract tests.
17
27
  - Never use unauthorized production personal data, credentials, or secrets in tests.
18
28
  - Do not hide failures by weakening assertions, deleting or skipping tests, excluding coverage, or
19
- adding indiscriminate retries. Change tests only when the intended behavior has changed.
29
+ adding indiscriminate retries. Correct or consolidate tests when their expectations are
30
+ wrong, obsolete, or redundant; preserve coverage of the accepted contract and follow project
31
+ policy for changes to mandatory gates.
20
32
  - If the affected scope cannot be established confidently, broaden the validation.
@@ -107,6 +107,24 @@ for consequential context gaps supported by repeated questions, repository evide
107
107
  or a concrete blocked task. Supplement only the actionable information needed to
108
108
  resolve the gap, using established references where possible.
109
109
 
110
+ Look for guard proliferation where checks duplicate assurance, constrain incidental
111
+ wording or implementation details, or repeatedly confirm another safeguard's
112
+ existence without protecting a consequential outcome. A check's subject, nesting
113
+ depth, or lack of recorded incidents is a review signal, not a removal rule.
114
+
115
+ For a candidate, identify the required outcome or constraint and the meaningful
116
+ coverage, independence, diagnosis, or recovery value it adds beyond existing
117
+ protection. Consider fixing the cause, repairing or consolidating existing checks,
118
+ replacing a mechanism, narrowing its scope, or retiring it rather than following
119
+ an escalation or downgrade ladder.
120
+
121
+ Compare the retained quality and credible risk with setup, execution, context,
122
+ false-alarm, review, and maintenance effort. Use qualitative judgment when it
123
+ settles the choice; measure or propose a bounded trial only for uncertainty that
124
+ could change the decision. State unsupported assumptions rather than inventing
125
+ incident probabilities or cost estimates. For consolidation or retirement,
126
+ explain how still-required protection remains covered, or why it is no longer needed.
127
+
110
128
  Choose **keep / rewrite / narrow / supplement / merge / relocate / retire / defer**.
111
129
  A useful tool can remain while its wrapper or routing changes. Equivalent content
112
130
  under the same scope can establish duplication without an incident or benchmark.
@@ -52,6 +52,21 @@ its assertions and conditions actually support each. Separate targets do not by
52
52
  themselves require separate runs or separate reviewers. Add checks for evidence gaps,
53
53
  credible risk, or applicable policy; required review independence is a separate issue.
54
54
 
55
+ For a new or materially changed executable safeguard, use proportionate evidence
56
+ that it accepts valid operation and detects the relevant violation. Existing tests,
57
+ controlled counterexamples, or another reliable method may suffice. A failing run
58
+ supports the claim only if it fails for the intended reason.
59
+
60
+ When a check can pass without exercising its intended scope, distinguish that
61
+ condition from a meaningful pass where the contract requires coverage. A known
62
+ violation can exercise detection without proving real-target discovery or required
63
+ execution.
64
+
65
+ Prefer existing test and execution mechanisms to a new supervisory layer.
66
+ Retain tests of safeguards when they close a meaningful evidence gap; adding a
67
+ check alone creates no need for another. Apply the evidence-reuse, invalidation,
68
+ and stopping guidance below.
69
+
55
70
  Identify binding tests and independent-review gates from their policy / CI source.
56
71
  Explicit project policy can bind even without CI enforcement. Preserve those gates;
57
72
  route disproportionate mandatory requirements to a separate policy decision. Retain
@@ -3,9 +3,10 @@ name: recurrence-prevention
3
3
  description: >-
4
4
  When the same defect, mistake, or incident happens AGAIN — a recurrence, not a one-off — verify
5
5
  it against prior evidence (memory, rule 근거 links, git/CHANGELOG history), classify it as a
6
- simple slip vs a complex harness problem, then escalate the countermeasure one level up the
7
- ladder: record (1st) → forced rule with linked priors (2nd) → structural gate — test, hook, or
8
- derive — once prose has failed (3rd+). Complex problems get countermeasure candidates designed
6
+ simple slip vs a complex harness problem, then re-evaluate the countermeasure that failed:
7
+ repair or replace it, choosing a record, a one-line criterion, a code fix, or a structural gate
8
+ (test, hook, derive) by evidence — more deterministic when prose has demonstrably failed,
9
+ nothing new when repair suffices. Complex problems get countermeasure candidates designed
9
10
  by a multi-persona panel instead of a quick patch. Use for "재발했어", "같은 실수 또 했네",
10
11
  "이거 저번에도 그랬잖아", "재발방지 대책 등록해줘", "재발방지 룰 만들어", "this happened again",
11
12
  "same bug as last time", "add a recurrence countermeasure", "postmortem this failure". Do NOT use
@@ -18,15 +19,18 @@ description: >-
18
19
  A defect that happens once is a bug. A defect that happens **twice is a countermeasure failure** —
19
20
  whatever was supposed to prevent the second occurrence (a mental note, a memory entry, a rule)
20
21
  demonstrably did not. So on a recurrence, the unit of work is not the fix (you already know the
21
- fix — you applied it last time). The unit of work is the **countermeasure**, and the core move is:
22
- **escalate it one level, because the current level just failed.**
22
+ fix — you applied it last time). The unit of work is the **countermeasure**: re-evaluate the cause
23
+ analysis and whatever was supposed to prevent this, repair or replace it, and reach for a more
24
+ deterministic mechanism only when repairing the existing one cannot suffice. **A recurrence is
25
+ evidence for that re-evaluation, not an order to add a level.**
23
26
 
24
27
  This skill codifies a practice proven in the harness repo that ships it: the no-false-ship rule
25
28
  was created only after the *third* false-ship incident; CHANGELOG drift survived a written
26
29
  convention for seven releases and stopped only when a test gate enforced it; comment warnings
27
30
  against hardcoded-list drift failed twice before "derive to a single source" became mandatory.
28
- The pattern is consistent: **each enforcement level fails in a characteristic way, and the answer
29
- is the next level up — not a louder version of the same level.**
31
+ The pattern is consistent: **each enforcement level fails in a characteristic way, and when one
32
+ has demonstrably failed the answer is usually a more deterministic mechanism — not a louder
33
+ version of the same level, and not a new mechanism when the existing one can be repaired.**
30
34
 
31
35
  ## When to use
32
36
 
@@ -89,7 +93,7 @@ For the capture checklist (what to write down about the failure so the next coun
89
93
  | Correct behavior | Known, agreed, undisputed | Disputed, unclear, or trade-off-laden |
90
94
  | Why it recurred | Wasn't followed: forgot, skipped, overlooked | The countermeasure itself was wrong/insufficient, or cause spans components |
91
95
  | Typical examples | Forgot a checklist step; committed a forbidden file; skipped a verification | A "verified" path that still shipped broken; drift between N surfaces; a gate that passes while the behavior fails |
92
- | Path | **Escalation ladder** (Step 3a) | **Multi-persona countermeasure design** (Step 3b) |
96
+ | Path | **Choose the countermeasure level** (Step 3a) | **Multi-persona countermeasure design** (Step 3b) |
93
97
 
94
98
  Two quick discriminators:
95
99
  - Could you write the corrective rule in one sentence right now, with confidence nobody would
@@ -103,21 +107,22 @@ When the classification itself is unclear, the cause taxonomy in
103
107
  policy · coordination · loop) usually decides it: a `logic`/`state` cause with agreed correct
104
108
  behavior is a slip; a `policy`/`coordination` cause is almost always the complex path.
105
109
 
106
- ## Step 3a — Simple slip: the escalation ladder
110
+ ## Step 3a — Simple slip: choose the countermeasure level
107
111
 
108
- Enter the ladder at the level matching the count. **Never enter above your count** — a gate for a
109
- first-time slip is gate inflation, and every gate is permanent maintenance + false-positive cost.
110
- (Two exceptions. Downward-honest: a **registered countermeasure that failed** counts that failure
111
- at its own level — a violated Level-1 rule at count 2 legitimately escalates to Level 2. And
112
- cost-driven: the Level 1 pre-flight below sends a count-2 slip straight to Level 2 when the wrong
113
- action is deterministically detectable, because a gate is the *cheaper* artifact, not the stronger
114
- one.)
112
+ The levels below are **options ordered by determinism, not a mandatory sequence**. The count is
113
+ evidence of how far the current countermeasure has failed; it does not select the level by itself.
114
+ Choose the least burdensome option that reliably prevents the cause: repair or replace the failed
115
+ countermeasure first; prefer a gate when prose has demonstrably failed or when the wrong action is
116
+ deterministically detectable (a gate is then the *cheaper* artifact, not the stronger one); and
117
+ let credible risk, not the count, justify protection — a concrete irreversible-damage path earns a
118
+ gate at count 0, while a first-time slip with no credible risk earns none (every gate is permanent
119
+ maintenance + false-positive cost).
115
120
 
116
- | Level | When | Countermeasure | Characteristic failure of this level |
121
+ | Level | Typical evidence (a signal, not a trigger) | Countermeasure | Characteristic failure of this level |
117
122
  |---|---|---|---|
118
123
  | **0 기록** | 1st occurrence | Fix + durable record: memory/lessons entry with **Why** it matters and **How to apply**, or an occurrence link on a related rule's 신설 근거 line if one already exists | Records don't steer — nothing re-reads them at the decision moment |
119
124
  | **1 룰 강제 등록** | 2nd occurrence (the record failed) | Register a forced rule on the project's **always-loaded steering surface**, using the template below — one-line principle + a one-line "신설 근거" with links to each occurrence. `.claude/rules/<name>.md` (Claude Code) or a rules section in `AGENTS.md` (other CLIs) — and **verify it actually loads**: if the always-loaded context (CLAUDE.md / AGENTS.md) doesn't already pull that location in, reference the rule from it. A rule file nothing loads is still Level 0 with extra steps. **Shape**: the *minimum condition* that prevents the failure — what must be true, not how to do it — stated in positive form; a prohibition only where the damage is irreversible | Prose can be skimmed, forgotten under context pressure, or rationalized around |
120
- | **2 구조적 게이트** | 3rd+ occurrence, **or** a registered countermeasure failed — bypassed *or* followed as designed yet insufficient | Deterministic enforcement that does not depend on the agent reading anything: a test gate that fails CI, a pre-action hook that blocks the command (where the CLI supports hooks), or **derive-to-single-source** so the drift is structurally impossible. **Shape**: one level up means *more deterministic*, not *more forbidden* — the gate checks a result and leaves the model's judgment where no check can express it | Gates that never demonstrably fire; gates so noisy they get bypassed |
125
+ | **2 구조적 게이트** | 3rd+ occurrence, **or** a registered countermeasure failed — bypassed *or* followed as designed yet insufficient | Deterministic enforcement that does not depend on the agent reading anything: a test gate that fails CI, a pre-action hook that blocks the command (where the CLI supports hooks), or **derive-to-single-source** so the drift is structurally impossible. **Shape**: a stronger countermeasure means *more deterministic*, not *more forbidden* — the gate checks a result and leaves the model's judgment where no check can express it | Gates that never demonstrably fire; gates so noisy they get bypassed |
121
126
 
122
127
  Load-bearing principle at Level 2: **comment warnings and doc reminders are not a blocking
123
128
  mechanism.** If the countermeasure's effect depends on someone (human or model) reading prose at
@@ -140,10 +145,10 @@ every install, forever**, while a gate costs CI time and **zero** standing conte
140
145
  questions in order and stop at the first that decides:
141
146
 
142
147
  1. **Can the wrong action be detected deterministically?** — a failing test, a hook that exits
143
- non-zero, a derive that deletes the duplicated list. If yes, **write the gate instead, even at
144
- count 2.** This is the one sanctioned way to enter above your count, and it is not gate
145
- inflation: you are not buying stronger enforcement than the count justifies, you are picking
146
- the cheaper artifact for the same enforcement. Code answers what code can answer.
148
+ non-zero, a derive that deletes the duplicated list. If yes, **write the gate instead of a
149
+ rule.** That is not gate inflation: you are not buying stronger enforcement than the risk
150
+ justifies, you are picking the cheaper artifact for the same enforcement. Code answers what
151
+ code can answer.
147
152
  2. **Would the rule change behaviour that would otherwise be wrong?** If the corrective principle
148
153
  is something a competent agent does anyway, or it restates a rule that already exists, it buys
149
154
  nothing and bills every session. Stay at Level 0.
@@ -185,7 +190,8 @@ irreversible damage or a protected gate; otherwise state the condition and leave
185
190
 
186
191
  1. 즉시 정정 보고 (무엇이 어떻게 위반이었는지 명시)
187
192
  2. 발생을 이슈·ADR 에 기록하고 durable memory 에 추가한 뒤, 위 근거 줄의 횟수와 링크를 올린다
188
- 3. 재위반이면 구조적 게이트(테스트/훅/derive)로 승격 — 프로즈는 이미 두 번 실패했다
193
+ 3. 재위반이면 원인 분석과 이 룰을 재평가한다 — 룰을 고쳐 충분하면 거기서 끝내고, 프로즈가 실제로
194
+ 실패한 것이면 결정론적 장치(테스트/훅/derive)로 대체한다
189
195
  ```
190
196
 
191
197
  The 근거 line is not decoration — it is the recurrence counter for the *next* occurrence, and it
@@ -210,10 +216,10 @@ for mechanics — run 3-5 personas in parallel, independently, then synthesize):
210
216
  Synthesize into 2-3 concrete countermeasure options with costs, and present them as a decision
211
217
  (recommendation first, ASIS→TOBE contrast — see `user-centered-explanation` if bundled). The
212
218
  chosen option still lands on the ladder: it becomes a record, a rule, or a gate — the panel decides
213
- *what* the countermeasure is, the ladder decides *how hard* it is enforced. The
214
- never-above-your-count guard governs slips with no failed countermeasure; here, a prior
215
- countermeasure that fired as designed and still failed already justifies landing one level above
216
- it (a failed Level-1 rule → a Level-2 gate is escalation, not inflation).
219
+ *what* the countermeasure is, the evidence decides *how deterministic* it must be. A prior
220
+ countermeasure that fired as designed and still failed is evidence that repairing it is not
221
+ enough, which justifies a more deterministic mechanism (a failed rule → a gate); it is not a
222
+ mandate to add one — replacing the failed mechanism is the default, stacking a new one on top is not.
217
223
 
218
224
  When choosing among the panel's options, compare them on prevention strength, false-positive rate,
219
225
  standing context cost, **per-run speed cost** (time, rounds, and questions added to every task),
@@ -232,7 +238,7 @@ would be a false ship. Before closing:
232
238
  test gate or derive.
233
239
  - **Derive/single-source**: show the duplicated site is gone (grep returns one definition).
234
240
  - **Rule (prose)**: not mechanically verifiable — say so explicitly ("등록됨, 준수는 미검증").
235
- That honesty is what justifies escalating to Level 2 if it recurs anyway.
241
+ That honesty is what justifies replacing it with a deterministic mechanism if it recurs anyway.
236
242
 
237
243
  ## Output format — 재발방지 보고
238
244
 
@@ -278,7 +284,8 @@ before creating them. Level 0 records need no confirmation.
278
284
  of prohibitions either routes around it or freezes. Write the state that must hold; reserve
279
285
  prohibitions for irreversible damage and protected gates.
280
286
  - **Same-level retry** — responding to a recurrence by rewriting the same rule more emphatically
281
- (CAPS, "NEVER", repetition). The level failed, not the wording. Escalate.
287
+ (CAPS, "NEVER", repetition). The level failed, not the wording — repair it if it can be
288
+ repaired, otherwise replace it with a more deterministic mechanism.
282
289
  - **Inflating the count** — when no prior artifact can be found, record this as the first confirmed
283
290
  occurrence rather than borrowing a remembered one to reach count 2.
284
291