@uzysjung/agent-harness 26.156.0 → 26.158.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.ko.md +66 -48
- package/README.md +64 -44
- package/dist/{chunk-PPJUFG53.js → chunk-HHCNPPGS.js} +5 -2
- package/dist/chunk-HHCNPPGS.js.map +1 -0
- package/dist/index.js +2 -2
- package/dist/index.js.map +1 -1
- package/dist/trust-tier-drift.js +1 -1
- package/package.json +1 -1
- package/templates/rules/test-policy.md +13 -1
- package/templates/skills/audit-harness-fit/references/audit.md +18 -0
- package/templates/skills/audit-harness-fit/references/verification.md +15 -0
- package/templates/skills/recurrence-prevention/SKILL.md +36 -29
- package/dist/chunk-PPJUFG53.js.map +0 -1
package/dist/trust-tier-drift.js
CHANGED
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@uzysjung/agent-harness",
|
|
3
|
-
"version": "26.
|
|
3
|
+
"version": "26.158.0",
|
|
4
4
|
"description": "Curate vetted AI-coding skills & plugins by your tech stack — install only what you need, across Claude Code, Codex, OpenCode & Antigravity",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"publishConfig": {
|
|
@@ -6,6 +6,16 @@ belongs to the Delivery rule.
|
|
|
6
6
|
- An ordinary change gets the repository's baseline CI and regression across the affected scope;
|
|
7
7
|
independent verification only where the Delivery rule requires it. A high-risk one widens that in
|
|
8
8
|
proportion to what it touches and what its failure would cost.
|
|
9
|
+
- Preserve required outcomes with the least burdensome reliable protection.
|
|
10
|
+
Prefer fixing the cause and reusing or improving existing mechanisms within scope;
|
|
11
|
+
no new guard is needed when those suffice. Judge safeguards by credible risk,
|
|
12
|
+
the assurance they add, and total effort, not request emphasis, incident counts,
|
|
13
|
+
or the number or subject of checks. Tests of safeguards follow the same standard.
|
|
14
|
+
Reuse valid evidence and recheck affected claims when relevant conditions change.
|
|
15
|
+
Within project policy, adapt methods to demonstrated model and tool capabilities
|
|
16
|
+
while preserving acceptance criteria, required gates, and authority boundaries.
|
|
17
|
+
Finish optional verification when required outcomes have sufficient evidence
|
|
18
|
+
and residual risk is within the project's accepted limits.
|
|
9
19
|
- High-risk includes at least authentication, authorization, payments and settlement, personal
|
|
10
20
|
data, data integrity, concurrency, state transitions, and migrations.
|
|
11
21
|
- Full regression, full E2E, full mutation, and periodic security scanning belong to the CI/CD
|
|
@@ -16,5 +26,7 @@ belongs to the Delivery rule.
|
|
|
16
26
|
the result. Otherwise, use explicit test doubles or contract tests.
|
|
17
27
|
- Never use unauthorized production personal data, credentials, or secrets in tests.
|
|
18
28
|
- Do not hide failures by weakening assertions, deleting or skipping tests, excluding coverage, or
|
|
19
|
-
adding indiscriminate retries.
|
|
29
|
+
adding indiscriminate retries. Correct or consolidate tests when their expectations are
|
|
30
|
+
wrong, obsolete, or redundant; preserve coverage of the accepted contract and follow project
|
|
31
|
+
policy for changes to mandatory gates.
|
|
20
32
|
- If the affected scope cannot be established confidently, broaden the validation.
|
|
@@ -107,6 +107,24 @@ for consequential context gaps supported by repeated questions, repository evide
|
|
|
107
107
|
or a concrete blocked task. Supplement only the actionable information needed to
|
|
108
108
|
resolve the gap, using established references where possible.
|
|
109
109
|
|
|
110
|
+
Look for guard proliferation where checks duplicate assurance, constrain incidental
|
|
111
|
+
wording or implementation details, or repeatedly confirm another safeguard's
|
|
112
|
+
existence without protecting a consequential outcome. A check's subject, nesting
|
|
113
|
+
depth, or lack of recorded incidents is a review signal, not a removal rule.
|
|
114
|
+
|
|
115
|
+
For a candidate, identify the required outcome or constraint and the meaningful
|
|
116
|
+
coverage, independence, diagnosis, or recovery value it adds beyond existing
|
|
117
|
+
protection. Consider fixing the cause, repairing or consolidating existing checks,
|
|
118
|
+
replacing a mechanism, narrowing its scope, or retiring it rather than following
|
|
119
|
+
an escalation or downgrade ladder.
|
|
120
|
+
|
|
121
|
+
Compare the retained quality and credible risk with setup, execution, context,
|
|
122
|
+
false-alarm, review, and maintenance effort. Use qualitative judgment when it
|
|
123
|
+
settles the choice; measure or propose a bounded trial only for uncertainty that
|
|
124
|
+
could change the decision. State unsupported assumptions rather than inventing
|
|
125
|
+
incident probabilities or cost estimates. For consolidation or retirement,
|
|
126
|
+
explain how still-required protection remains covered, or why it is no longer needed.
|
|
127
|
+
|
|
110
128
|
Choose **keep / rewrite / narrow / supplement / merge / relocate / retire / defer**.
|
|
111
129
|
A useful tool can remain while its wrapper or routing changes. Equivalent content
|
|
112
130
|
under the same scope can establish duplication without an incident or benchmark.
|
|
@@ -52,6 +52,21 @@ its assertions and conditions actually support each. Separate targets do not by
|
|
|
52
52
|
themselves require separate runs or separate reviewers. Add checks for evidence gaps,
|
|
53
53
|
credible risk, or applicable policy; required review independence is a separate issue.
|
|
54
54
|
|
|
55
|
+
For a new or materially changed executable safeguard, use proportionate evidence
|
|
56
|
+
that it accepts valid operation and detects the relevant violation. Existing tests,
|
|
57
|
+
controlled counterexamples, or another reliable method may suffice. A failing run
|
|
58
|
+
supports the claim only if it fails for the intended reason.
|
|
59
|
+
|
|
60
|
+
When a check can pass without exercising its intended scope, distinguish that
|
|
61
|
+
condition from a meaningful pass where the contract requires coverage. A known
|
|
62
|
+
violation can exercise detection without proving real-target discovery or required
|
|
63
|
+
execution.
|
|
64
|
+
|
|
65
|
+
Prefer existing test and execution mechanisms to a new supervisory layer.
|
|
66
|
+
Retain tests of safeguards when they close a meaningful evidence gap; adding a
|
|
67
|
+
check alone creates no need for another. Apply the evidence-reuse, invalidation,
|
|
68
|
+
and stopping guidance below.
|
|
69
|
+
|
|
55
70
|
Identify binding tests and independent-review gates from their policy / CI source.
|
|
56
71
|
Explicit project policy can bind even without CI enforcement. Preserve those gates;
|
|
57
72
|
route disproportionate mandatory requirements to a separate policy decision. Retain
|
|
@@ -3,9 +3,10 @@ name: recurrence-prevention
|
|
|
3
3
|
description: >-
|
|
4
4
|
When the same defect, mistake, or incident happens AGAIN — a recurrence, not a one-off — verify
|
|
5
5
|
it against prior evidence (memory, rule 근거 links, git/CHANGELOG history), classify it as a
|
|
6
|
-
simple slip vs a complex harness problem, then
|
|
7
|
-
|
|
8
|
-
|
|
6
|
+
simple slip vs a complex harness problem, then re-evaluate the countermeasure that failed:
|
|
7
|
+
repair or replace it, choosing a record, a one-line criterion, a code fix, or a structural gate
|
|
8
|
+
(test, hook, derive) by evidence — more deterministic when prose has demonstrably failed,
|
|
9
|
+
nothing new when repair suffices. Complex problems get countermeasure candidates designed
|
|
9
10
|
by a multi-persona panel instead of a quick patch. Use for "재발했어", "같은 실수 또 했네",
|
|
10
11
|
"이거 저번에도 그랬잖아", "재발방지 대책 등록해줘", "재발방지 룰 만들어", "this happened again",
|
|
11
12
|
"same bug as last time", "add a recurrence countermeasure", "postmortem this failure". Do NOT use
|
|
@@ -18,15 +19,18 @@ description: >-
|
|
|
18
19
|
A defect that happens once is a bug. A defect that happens **twice is a countermeasure failure** —
|
|
19
20
|
whatever was supposed to prevent the second occurrence (a mental note, a memory entry, a rule)
|
|
20
21
|
demonstrably did not. So on a recurrence, the unit of work is not the fix (you already know the
|
|
21
|
-
fix — you applied it last time). The unit of work is the **countermeasure
|
|
22
|
-
|
|
22
|
+
fix — you applied it last time). The unit of work is the **countermeasure**: re-evaluate the cause
|
|
23
|
+
analysis and whatever was supposed to prevent this, repair or replace it, and reach for a more
|
|
24
|
+
deterministic mechanism only when repairing the existing one cannot suffice. **A recurrence is
|
|
25
|
+
evidence for that re-evaluation, not an order to add a level.**
|
|
23
26
|
|
|
24
27
|
This skill codifies a practice proven in the harness repo that ships it: the no-false-ship rule
|
|
25
28
|
was created only after the *third* false-ship incident; CHANGELOG drift survived a written
|
|
26
29
|
convention for seven releases and stopped only when a test gate enforced it; comment warnings
|
|
27
30
|
against hardcoded-list drift failed twice before "derive to a single source" became mandatory.
|
|
28
|
-
The pattern is consistent: **each enforcement level fails in a characteristic way, and
|
|
29
|
-
|
|
31
|
+
The pattern is consistent: **each enforcement level fails in a characteristic way, and when one
|
|
32
|
+
has demonstrably failed the answer is usually a more deterministic mechanism — not a louder
|
|
33
|
+
version of the same level, and not a new mechanism when the existing one can be repaired.**
|
|
30
34
|
|
|
31
35
|
## When to use
|
|
32
36
|
|
|
@@ -89,7 +93,7 @@ For the capture checklist (what to write down about the failure so the next coun
|
|
|
89
93
|
| Correct behavior | Known, agreed, undisputed | Disputed, unclear, or trade-off-laden |
|
|
90
94
|
| Why it recurred | Wasn't followed: forgot, skipped, overlooked | The countermeasure itself was wrong/insufficient, or cause spans components |
|
|
91
95
|
| Typical examples | Forgot a checklist step; committed a forbidden file; skipped a verification | A "verified" path that still shipped broken; drift between N surfaces; a gate that passes while the behavior fails |
|
|
92
|
-
| Path | **
|
|
96
|
+
| Path | **Choose the countermeasure level** (Step 3a) | **Multi-persona countermeasure design** (Step 3b) |
|
|
93
97
|
|
|
94
98
|
Two quick discriminators:
|
|
95
99
|
- Could you write the corrective rule in one sentence right now, with confidence nobody would
|
|
@@ -103,21 +107,22 @@ When the classification itself is unclear, the cause taxonomy in
|
|
|
103
107
|
policy · coordination · loop) usually decides it: a `logic`/`state` cause with agreed correct
|
|
104
108
|
behavior is a slip; a `policy`/`coordination` cause is almost always the complex path.
|
|
105
109
|
|
|
106
|
-
## Step 3a — Simple slip: the
|
|
110
|
+
## Step 3a — Simple slip: choose the countermeasure level
|
|
107
111
|
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
112
|
+
The levels below are **options ordered by determinism, not a mandatory sequence**. The count is
|
|
113
|
+
evidence of how far the current countermeasure has failed; it does not select the level by itself.
|
|
114
|
+
Choose the least burdensome option that reliably prevents the cause: repair or replace the failed
|
|
115
|
+
countermeasure first; prefer a gate when prose has demonstrably failed or when the wrong action is
|
|
116
|
+
deterministically detectable (a gate is then the *cheaper* artifact, not the stronger one); and
|
|
117
|
+
let credible risk, not the count, justify protection — a concrete irreversible-damage path earns a
|
|
118
|
+
gate at count 0, while a first-time slip with no credible risk earns none (every gate is permanent
|
|
119
|
+
maintenance + false-positive cost).
|
|
115
120
|
|
|
116
|
-
| Level |
|
|
121
|
+
| Level | Typical evidence (a signal, not a trigger) | Countermeasure | Characteristic failure of this level |
|
|
117
122
|
|---|---|---|---|
|
|
118
123
|
| **0 기록** | 1st occurrence | Fix + durable record: memory/lessons entry with **Why** it matters and **How to apply**, or an occurrence link on a related rule's 신설 근거 line if one already exists | Records don't steer — nothing re-reads them at the decision moment |
|
|
119
124
|
| **1 룰 강제 등록** | 2nd occurrence (the record failed) | Register a forced rule on the project's **always-loaded steering surface**, using the template below — one-line principle + a one-line "신설 근거" with links to each occurrence. `.claude/rules/<name>.md` (Claude Code) or a rules section in `AGENTS.md` (other CLIs) — and **verify it actually loads**: if the always-loaded context (CLAUDE.md / AGENTS.md) doesn't already pull that location in, reference the rule from it. A rule file nothing loads is still Level 0 with extra steps. **Shape**: the *minimum condition* that prevents the failure — what must be true, not how to do it — stated in positive form; a prohibition only where the damage is irreversible | Prose can be skimmed, forgotten under context pressure, or rationalized around |
|
|
120
|
-
| **2 구조적 게이트** | 3rd+ occurrence, **or** a registered countermeasure failed — bypassed *or* followed as designed yet insufficient | Deterministic enforcement that does not depend on the agent reading anything: a test gate that fails CI, a pre-action hook that blocks the command (where the CLI supports hooks), or **derive-to-single-source** so the drift is structurally impossible. **Shape**:
|
|
125
|
+
| **2 구조적 게이트** | 3rd+ occurrence, **or** a registered countermeasure failed — bypassed *or* followed as designed yet insufficient | Deterministic enforcement that does not depend on the agent reading anything: a test gate that fails CI, a pre-action hook that blocks the command (where the CLI supports hooks), or **derive-to-single-source** so the drift is structurally impossible. **Shape**: a stronger countermeasure means *more deterministic*, not *more forbidden* — the gate checks a result and leaves the model's judgment where no check can express it | Gates that never demonstrably fire; gates so noisy they get bypassed |
|
|
121
126
|
|
|
122
127
|
Load-bearing principle at Level 2: **comment warnings and doc reminders are not a blocking
|
|
123
128
|
mechanism.** If the countermeasure's effect depends on someone (human or model) reading prose at
|
|
@@ -140,10 +145,10 @@ every install, forever**, while a gate costs CI time and **zero** standing conte
|
|
|
140
145
|
questions in order and stop at the first that decides:
|
|
141
146
|
|
|
142
147
|
1. **Can the wrong action be detected deterministically?** — a failing test, a hook that exits
|
|
143
|
-
non-zero, a derive that deletes the duplicated list. If yes, **write the gate instead
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
|
|
148
|
+
non-zero, a derive that deletes the duplicated list. If yes, **write the gate instead of a
|
|
149
|
+
rule.** That is not gate inflation: you are not buying stronger enforcement than the risk
|
|
150
|
+
justifies, you are picking the cheaper artifact for the same enforcement. Code answers what
|
|
151
|
+
code can answer.
|
|
147
152
|
2. **Would the rule change behaviour that would otherwise be wrong?** If the corrective principle
|
|
148
153
|
is something a competent agent does anyway, or it restates a rule that already exists, it buys
|
|
149
154
|
nothing and bills every session. Stay at Level 0.
|
|
@@ -185,7 +190,8 @@ irreversible damage or a protected gate; otherwise state the condition and leave
|
|
|
185
190
|
|
|
186
191
|
1. 즉시 정정 보고 (무엇이 어떻게 위반이었는지 명시)
|
|
187
192
|
2. 발생을 이슈·ADR 에 기록하고 durable memory 에 추가한 뒤, 위 근거 줄의 횟수와 링크를 올린다
|
|
188
|
-
3. 재위반이면
|
|
193
|
+
3. 재위반이면 원인 분석과 이 룰을 재평가한다 — 룰을 고쳐 충분하면 거기서 끝내고, 프로즈가 실제로
|
|
194
|
+
실패한 것이면 결정론적 장치(테스트/훅/derive)로 대체한다
|
|
189
195
|
```
|
|
190
196
|
|
|
191
197
|
The 근거 line is not decoration — it is the recurrence counter for the *next* occurrence, and it
|
|
@@ -210,10 +216,10 @@ for mechanics — run 3-5 personas in parallel, independently, then synthesize):
|
|
|
210
216
|
Synthesize into 2-3 concrete countermeasure options with costs, and present them as a decision
|
|
211
217
|
(recommendation first, ASIS→TOBE contrast — see `user-centered-explanation` if bundled). The
|
|
212
218
|
chosen option still lands on the ladder: it becomes a record, a rule, or a gate — the panel decides
|
|
213
|
-
*what* the countermeasure is, the
|
|
214
|
-
|
|
215
|
-
|
|
216
|
-
|
|
219
|
+
*what* the countermeasure is, the evidence decides *how deterministic* it must be. A prior
|
|
220
|
+
countermeasure that fired as designed and still failed is evidence that repairing it is not
|
|
221
|
+
enough, which justifies a more deterministic mechanism (a failed rule → a gate); it is not a
|
|
222
|
+
mandate to add one — replacing the failed mechanism is the default, stacking a new one on top is not.
|
|
217
223
|
|
|
218
224
|
When choosing among the panel's options, compare them on prevention strength, false-positive rate,
|
|
219
225
|
standing context cost, **per-run speed cost** (time, rounds, and questions added to every task),
|
|
@@ -232,7 +238,7 @@ would be a false ship. Before closing:
|
|
|
232
238
|
test gate or derive.
|
|
233
239
|
- **Derive/single-source**: show the duplicated site is gone (grep returns one definition).
|
|
234
240
|
- **Rule (prose)**: not mechanically verifiable — say so explicitly ("등록됨, 준수는 미검증").
|
|
235
|
-
That honesty is what justifies
|
|
241
|
+
That honesty is what justifies replacing it with a deterministic mechanism if it recurs anyway.
|
|
236
242
|
|
|
237
243
|
## Output format — 재발방지 보고
|
|
238
244
|
|
|
@@ -278,7 +284,8 @@ before creating them. Level 0 records need no confirmation.
|
|
|
278
284
|
of prohibitions either routes around it or freezes. Write the state that must hold; reserve
|
|
279
285
|
prohibitions for irreversible damage and protected gates.
|
|
280
286
|
- **Same-level retry** — responding to a recurrence by rewriting the same rule more emphatically
|
|
281
|
-
(CAPS, "NEVER", repetition). The level failed, not the wording
|
|
287
|
+
(CAPS, "NEVER", repetition). The level failed, not the wording — repair it if it can be
|
|
288
|
+
repaired, otherwise replace it with a more deterministic mechanism.
|
|
282
289
|
- **Inflating the count** — when no prior artifact can be found, record this as the first confirmed
|
|
283
290
|
occurrence rather than borrowing a remembered one to reach count 2.
|
|
284
291
|
|