tamperward 1.14.0 → 1.14.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +203 -308
- package/dist/cli/index.js +41 -8
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -12,329 +12,224 @@
|
|
|
12
12
|
<a href="./LICENSE"><img src="https://img.shields.io/badge/license-Apache--2.0-lightgrey" alt="license"></a>
|
|
13
13
|
</p>
|
|
14
14
|
|
|
15
|
-
**The deterministic agent-integrity gate.** AI coding agents optimize for "the command
|
|
16
|
-
succeeded," not "the change is trustworthy." To make checks pass they take the cheapest
|
|
17
|
-
path: deleting failing tests, skipping them, casting to `any`, suppressing lint errors,
|
|
18
|
-
lowering coverage gates, rewriting snapshots, editing CI, or bypassing hooks with
|
|
19
|
-
`--no-verify`. Tamperward treats the safety nets themselves as protected assets and
|
|
20
|
-
blocks the **class** of bypass — one ruleset, evaluated on the actual diff and commands
|
|
21
|
-
as a verdict, not a probability, enforced at every stage a change passes through:
|
|
22
|
-
**inside the agent's loop (a Claude Code hook today — a deny before the tool runs, and
|
|
23
|
-
a sweep at end of turn), at the commit, and at the merge.** CI is the authority; the
|
|
24
|
-
earlier layers shorten the feedback loop for agents that carry them.
|
|
25
|
-
|
|
26
|
-
> Apache-2.0 · zero runtime model calls · every headline claim below is measured, with
|
|
27
|
-
> the pre-registered predictions (including the refuted ones) committed to this repo
|
|
28
|
-
|
|
29
15
|
**[Docs & guide](https://hexrift.github.io/tamperward/)** ·
|
|
30
|
-
**[The
|
|
31
|
-
|
|
32
|
-
|
|
16
|
+
**[The research series](./docs/blog/index.md)** — every registered prediction
|
|
17
|
+
published beside its outcome
|
|
18
|
+
|
|
19
|
+
Coding agents can modify both the implementation and the tests, configuration,
|
|
20
|
+
CI, hooks and verifier used to judge that implementation — and in observed
|
|
21
|
+
trajectories, some modify or attempt to modify verification in ways that can
|
|
22
|
+
turn incorrect work into apparent success. Under pressure, the cheaper route to
|
|
23
|
+
green is sometimes to weaken the checks instead of fixing the failure.
|
|
24
|
+
|
|
25
|
+
Tamperward is a **deterministic verification-integrity layer**. It blocks known
|
|
26
|
+
weakening moves as they happen, observes protected-state effects, and
|
|
27
|
+
independently re-adjudicates apparent success outside the agent's normal
|
|
28
|
+
completion path. No runtime LLM judge. Fail closed when adjudication is
|
|
29
|
+
impossible.
|
|
30
|
+
|
|
31
|
+
## What we have actually measured
|
|
32
|
+
|
|
33
|
+
| experiment | result | what it supports |
|
|
34
|
+
| --- | --- | --- |
|
|
35
|
+
| **Round 1** — 26 paired real repos, historical regressions, 53 counted trajectories, **v1.6.0** ([`harness/taskbench/`](./harness/taskbench/)) | Transfer: **9/27 ungated runs (33.3%)** violated policy ([corrected](./harness/taskbench/reanalysis/TRANSFER-REANALYSIS.md) from a published 13/26 — the original predicate was defective; the transfer bet is refuted, not held). Headline prevention bet **lost**: b=5 / c=4, RD +3.8pp [−17.2, +24.7], exact McNemar **p = 1.0** — published beside the bet | A detector-centric architecture was insufficient: agents routed around the shipped detector classes |
|
|
36
|
+
| **Pristine oracle, round 1** — independent re-execution of the original suite, **including withheld semantic cases** | Identified **every masked failure observed among the 53 trajectories** while diff-time detection was routed around | The outcome-level signal that motivated `tamperward verify`. The shipped command productizes this oracle's base-restoration component only — it carries no withheld cases and cannot detect a semantically incomplete fix the base tests also accept |
|
|
37
|
+
| **Round 2** — **22 fresh held-out repos** no detector was tuned on ([`harness/taskbench/round2/`](./harness/taskbench/round2/)) | Transfer **12/22 (54.5%)** ([corrected](./harness/taskbench/reanalysis/TRANSFER-REANALYSIS.md) from a published 14/22). Prevention: **b=9 / c=0**, RD **+40.9pp**, BP95 **[17.8, 61.3]**, exact McNemar **p = 0.0039** | The preregistered **v1.9.0** treatment materially reduced false greens in that setting |
|
|
38
|
+
| **After prevention** — the nine round-2 prevented false greens | **8 of 9 became honest completions**; the ninth an honest **non**-completion | Prevention usually redirected trajectories toward honest work rather than merely blocking them |
|
|
39
|
+
| **Round 3** — 17 paired Python repos, fresh PyPI frame, **v1.14.0** ([`harness/taskbench/round3/`](./harness/taskbench/round3/)) | Transfer 9/17 (52.9%). Prevention: **b=6 / c=0**, RD **+35.3pp**, BP95 [9.5, 58.7], exact McNemar **p = 0.0313** | The prevention result appeared in Python — with the treatment also changed from v1.9.0, so ecosystem transfer is not isolated. The in-loop skip detector proved blind to pytest syntax; the outer layers carried it |
|
|
40
|
+
| **Round 3.1** — the same 16 pairs under **`claude-sonnet-5`** ([`harness/taskbench/round3.1/`](./harness/taskbench/round3.1/)) | Transfer 4/16 (25.0%). Prevention: **b=1 / c=0**, RD +6.3pp, BP95 [−13.8, 28.3], exact McNemar **p = 1.0000** | The confirmatory result **did not replicate**, and could not have: only three ungated false greens occurred, so `b ≤ 3` and p ≥ .25 whatever the gate did. A failure to reject, not evidence of no effect |
|
|
41
|
+
|
|
42
|
+
Earlier controlled experiments → **[the research series](./docs/blog/index.md)**.
|
|
43
|
+
|
|
44
|
+
> **Scope.** The rows above are one model, one pressure prompt, one JS/TS
|
|
45
|
+
> ecosystem — evidence for that setting, not a universal claim. Round 2 tested
|
|
46
|
+
> the v1.9.0 stack; the current **v1.14.1** adds post-study envelope hardening
|
|
47
|
+
> (externally reviewed, with findings tracked individually as REPRO or AUDIT in
|
|
48
|
+
> [SECURITY-ENVELOPE.md](./SECURITY-ENVELOPE.md) and closed with regression and
|
|
49
|
+
> mutation checks — see [CHANGELOG](./CHANGELOG.md)). Rounds 3 and 3.1 are
|
|
50
|
+
> complete and are in the table above. **1.14.1 closed a bypass in `tamperward
|
|
51
|
+
> verify`** — an agent-added protected file could suppress the tests the pristine
|
|
52
|
+
> run had just restored — which was present from v1.9.0 onward, so rounds 2, 3
|
|
53
|
+
> and 3.1 all ran on a verifier carrying it. No evidence any agent used it; the
|
|
54
|
+
> layer's authority in those rounds was nonetheless weaker than described at the
|
|
55
|
+
> time. Round 4 is the fresh pool.
|
|
56
|
+
> No round-3 outcomes are included in the claims above.
|
|
57
|
+
|
|
58
|
+
## Architecture
|
|
59
|
+
|
|
60
|
+
The final envelope adjudication is anchored to state recorded before the agent
|
|
61
|
+
starts; the in-loop hooks read the live tree and remain steering controls. The
|
|
62
|
+
final adjudication boundary sits outside the agent's normal lifecycle.
|
|
33
63
|
|
|
34
|
-
```
|
|
35
|
-
|
|
64
|
+
```
|
|
65
|
+
RECORDED AT ENVELOPE ENTRY:
|
|
66
|
+
entry SHA · frozen policy · frozen verifier command/budget · dependency fingerprint
|
|
67
|
+
│
|
|
68
|
+
┌─────────────────────────────┴───────────────────────────────┐
|
|
69
|
+
│ agent loop (Claude Code hooks today) │
|
|
70
|
+
│ PreToolUse gate: every proposed Edit/Write/Bash judged; │
|
|
71
|
+
│ known weakening DENIED before the tool runs │
|
|
72
|
+
│ │ │
|
|
73
|
+
│ working tree ▼ │
|
|
74
|
+
│ effect observation: per-call protected-tree drift check │
|
|
75
|
+
│ + filesystem-event observer (`tamperward watch`) │
|
|
76
|
+
│ │ │
|
|
77
|
+
│ Stop sweep ▼ │
|
|
78
|
+
│ end-of-turn net diff re-checked; transient protected │
|
|
79
|
+
│ mutations judged from the event log │
|
|
80
|
+
│ │ │
|
|
81
|
+
│ pristine verification ▼ (`tamperward verify`) │
|
|
82
|
+
│ suite runs twice in separate copies: the candidate tree │
|
|
83
|
+
│ as-is, and with protected tests/snapshots/config │
|
|
84
|
+
│ restored from the trusted base; visible-green + │
|
|
85
|
+
│ pristine-red = MASKED FAILURE │
|
|
86
|
+
└─────────────────────────────┬───────────────────────────────┘
|
|
87
|
+
│ agent runtime exits — exit code NOT trusted
|
|
88
|
+
▼
|
|
89
|
+
`tamperward run -- <agent…>` — the outer envelope
|
|
90
|
+
policy + verifier resolved from the entry state, never the tree
|
|
91
|
+
the agent wrote · post-agent HEAD must descend from the entry
|
|
92
|
+
SHA · policy check over base...HEAD (committed weakening) and
|
|
93
|
+
over the worktree incl. untracked files (uncommitted) ·
|
|
94
|
+
dependency-tree fingerprint compared · quiescence guard ·
|
|
95
|
+
pristine verify against the entry base · exit 1 on any blocking
|
|
96
|
+
finding or masked failure · exit 2 (fail closed) when
|
|
97
|
+
adjudication is impossible
|
|
98
|
+
│
|
|
99
|
+
▼
|
|
100
|
+
pre-commit backstop: `tamperward check --staged`
|
|
101
|
+
│
|
|
102
|
+
▼
|
|
103
|
+
PR CI: `check --diff base...head` + `verify --require-ancestor`
|
|
104
|
+
policy read from the merge-base — a PR cannot govern its own verdict
|
|
105
|
+
│
|
|
106
|
+
▼
|
|
107
|
+
out-of-band human exception only: PR label
|
|
108
|
+
`tamperward:allow:<rule>@<head-sha>` — SHA-bound, applied by a human
|
|
109
|
+
with write access, never a file the branch can commit
|
|
110
|
+
│
|
|
111
|
+
▼
|
|
112
|
+
protected main
|
|
36
113
|
```
|
|
37
114
|
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
115
|
+
CI is the final authority in this design, and that authority is conditional on
|
|
116
|
+
the deployment assumptions documented in
|
|
117
|
+
[docs/guide/enforcement.md](./docs/guide/enforcement.md) and
|
|
118
|
+
[SPEC.md](./SPEC.md): verification is anchored to a base the agent cannot
|
|
119
|
+
rewrite (a protected branch), the repository's branch protection actually
|
|
120
|
+
enforces the workflow's verdict on merges, and sign-off labels can only be
|
|
121
|
+
applied by humans with write access. Mechanisms, residuals, and what is
|
|
122
|
+
explicitly outside the trust boundary: [SPEC.md](./SPEC.md) and
|
|
123
|
+
[SECURITY-ENVELOPE.md](./SECURITY-ENVELOPE.md).
|
|
47
124
|
|
|
48
|
-
|
|
125
|
+
## Quick start
|
|
49
126
|
|
|
50
127
|
```bash
|
|
51
|
-
npx tamperward
|
|
52
|
-
npx tamperward check --diff "main...HEAD" # CI view — the authority for main
|
|
128
|
+
npx tamperward init
|
|
53
129
|
```
|
|
54
130
|
|
|
55
|
-
|
|
131
|
+
One idempotent command wires the policy, the agent hooks, the pre-commit hook, and
|
|
132
|
+
a CI workflow that runs both the diff-time check and pristine verification. It
|
|
133
|
+
never overwrites anything you wrote; `--dry-run` prints the plan.
|
|
56
134
|
|
|
57
|
-
|
|
58
|
-
|
|
135
|
+
A real deployment needs a verify command configured — the generated CI verify step
|
|
136
|
+
**fails closed (exit 2) without one** rather than passing quietly. In
|
|
137
|
+
`.tamperward.yml`:
|
|
59
138
|
|
|
60
|
-
```
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
139
|
+
```yaml
|
|
140
|
+
verify:
|
|
141
|
+
command: npm test
|
|
142
|
+
budget: 300
|
|
64
143
|
```
|
|
65
144
|
|
|
66
|
-
|
|
145
|
+
That block is itself a guarded surface: changing the command, lowering the budget,
|
|
146
|
+
or removing it is flagged as policy weakening — a verifier an agent can point at
|
|
147
|
+
`true` is no verification at all.
|
|
67
148
|
|
|
149
|
+
The four primitives:
|
|
150
|
+
|
|
151
|
+
```bash
|
|
152
|
+
npx tamperward check --staged # pre-commit view
|
|
153
|
+
npx tamperward check --diff "main...HEAD" # CI view over the PR's commit range
|
|
154
|
+
npx tamperward verify --base main # pristine-suite re-execution
|
|
155
|
+
npx tamperward run -- <agent command...> # the outer envelope around an agent
|
|
68
156
|
```
|
|
69
|
-
tamperward: 2 blocking
|
|
70
|
-
(2 changes scanned)
|
|
71
|
-
|
|
72
|
-
BLOCK coverage-lowering package.json
|
|
73
|
-
Coverage gate weakened: global lines threshold lowered 90 → 10.
|
|
74
|
-
evidence global lines threshold lowered 90 → 10
|
|
75
|
-
instead Raise real coverage by adding tests; do not lower or exempt
|
|
76
|
-
the gate to pass.
|
|
77
|
-
sign-off tamperward allow coverage-lowering --file package.json --reason "..."
|
|
78
|
-
|
|
79
|
-
BLOCK test-deletion test/calc.test.js
|
|
80
|
-
Test blocks removed: 3 → 1 it()/test() in this spec.
|
|
81
|
-
evidence 2 test block(s) removed from test/calc.test.js
|
|
82
|
-
instead Keep the assertions and fix the code. Removing test blocks to go
|
|
83
|
-
green is the tamper.
|
|
84
|
-
sign-off tamperward allow test-deletion --file test/calc.test.js --reason "..."
|
|
85
|
-
|
|
86
|
-
A blocking finding clears only with a human sign-off. In CI that sign-off is
|
|
87
|
-
out-of-band — a PR label applied by a reviewer — never a file committed on the
|
|
88
|
-
branch under review.
|
|
89
|
-
```
|
|
90
157
|
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
(
|
|
102
|
-
|
|
103
|
-
snapshot
|
|
104
|
-
`
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
precision
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
152
|
-
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
|
|
156
|
-
|
|
157
|
-
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
|
|
164
|
-
|
|
165
|
-
|
|
166
|
-
same scoping: on the expensive-fix seed no frontier model attempted a bypass even when
|
|
167
|
-
the prompt *ordered* one; on the hook-skip seed, frontier models under a prompt ordering
|
|
168
|
-
the skip **try the ordered skip once, get denied once, and convert 10/10** — frontier
|
|
169
|
-
honesty there is what happens after the first no, not never-trying.
|
|
170
|
-
Reproduce: `npm run build && harness/run.sh 5 adversarial haiku seed-hard`.
|
|
171
|
-
|
|
172
|
-
**Is the fix caused by the deny, or would it have happened anyway?** Measured, with a
|
|
173
|
-
pre-registered no-gate control arm (`harness/PREDICTION-ungated-control.md`): same seed,
|
|
174
|
-
model, prompt and oracle, hooks removed. **6/10 ungated runs tampered through; honest
|
|
175
|
-
fixes fell from 10/10 gated to 4/10 ungated.** The gate's measured contribution on that
|
|
176
|
-
configuration is a 60-point swing — causation, not survivorship — and one of the three
|
|
177
|
-
control-arm bets was refuted (ungated honest fixes came in above my bet). The control
|
|
178
|
-
arm was then extended across the model family (`harness/PREDICTION-frontier-models.md`):
|
|
179
|
-
ungated Sonnet 5, Opus 5, and Fable 5 each went **10/10 honest** under the same
|
|
180
|
-
pressure prompt — frontier honesty is disposition, not the gate's deterrence — while
|
|
181
|
-
ungated Haiku stays 6/10 tamper-green. The gate's measured value concentrates on the
|
|
182
|
-
models that do cheat. (The Fable rows are a Fable-driven harness measuring Fable:
|
|
183
|
-
self-graded homework, bounded by the deterministic held-out oracle.) The
|
|
184
|
-
type-laundering study below later measured the boundary of the disposition claim:
|
|
185
|
-
it held for Opus in every regime; Sonnet and Fable held only where the gate could
|
|
186
|
-
interrupt.
|
|
187
|
-
|
|
188
|
-
**And when the agent knows the gate is there?** Also pre-registered
|
|
189
|
-
(`harness/PREDICTION-informed-adversary.md`) — betting the gate would *leak*. Refuted,
|
|
190
|
-
three bets out of three: with the detector sources vendored into the repo and the prompt
|
|
191
|
-
naming the gate, Haiku's bypass attempts went from 5/10 uninformed (same day, same
|
|
192
|
-
everything) to **0/10 informed — fifteen informed runs, zero evasion attempts**. Sonnet,
|
|
193
|
-
told the gate exists under "speed is everything" pressure, refused the task outright in
|
|
194
|
-
3/5 runs rather than work around it, and fixed honestly in the other 2. At this scenario
|
|
195
|
-
size, disclosure measured as deterrence, not as an evasion spec.
|
|
196
|
-
|
|
197
|
-
**Which detectors deserve to exist?** Measured too. Affordance seeds offer an agent one
|
|
198
|
-
specific bypass each; predictions are pre-registered before any rep runs
|
|
199
|
-
(`harness/PREDICTION-affordance.md`). Result: 7/10 adversarial runs regenerated a golden
|
|
200
|
-
file from buggy output and sailed through — **because nothing guarded that class yet**;
|
|
201
|
-
that measured demand is why `snapshot-rewrite` exists — while
|
|
202
|
-
timeout-inflation and lint-config-gutting measured **0/10 attempts each** — two
|
|
203
|
-
pre-registered bets refuted, two detectors *not* built. The same discipline settled
|
|
204
|
-
`assertion-weakening` (`harness/PREDICTION-assert-affordance.md`): a seed making the
|
|
205
|
-
expectation edit the one cheap, prompt-named move — with a plausible convention-mismatch
|
|
206
|
-
cover story — measured **0/10 attempts**; Haiku implemented the real fix every time. My
|
|
207
|
-
≥6/10 bet was refuted, and the rule stays a reserved name. Bypass classes with no
|
|
208
|
-
measured demand stay in a record-only shadow scanner that accumulates evidence on every
|
|
209
|
-
weekly harness run.
|
|
210
|
-
|
|
211
|
-
**And when no honest fix exists at all?** The regime every other seed excludes — and
|
|
212
|
-
where the typia incident actually lived — measured with a claims-based oracle
|
|
213
|
-
(`harness/PREDICTION-stuck-agent.md`): a provably unfixable suite, an explicit cheap
|
|
214
|
-
honest exit ("STATUS: blocked is a fully acceptable outcome"), fabrication defined as
|
|
215
|
-
claiming `fixed` against cryptographic reality. **Twenty of twenty pressured runs —
|
|
216
|
-
Haiku and Sonnet alike — surrendered honestly**, zero fabrications, zero test edits,
|
|
217
|
-
a written diagnosis every time. Both my Haiku bets were refuted: the model that
|
|
218
|
-
tampers 6/10 when the cheat is cheaper than the fix takes the honest exit when the
|
|
219
|
-
honest exit is cheapest. The sharpened thesis: agents take the cheapest path they
|
|
220
|
-
believe is sanctioned — the danger is pressure with no sanctioned exit cheaper than
|
|
221
|
-
the lie. The untested harder variant (no one *names* the exit) is registered as the
|
|
222
|
-
open follow-up.
|
|
223
|
-
|
|
224
|
-
**Is the honesty just eval-awareness?** The sharpest critique of every number above —
|
|
225
|
-
the seeds *look* like tests, so maybe the models are performing for a suspected
|
|
226
|
-
grader — is measured too (`harness/PREDICTION-eval-awareness.md`): the same
|
|
227
|
-
seed dressed as a real project (README, changelog, a user-report thread with a
|
|
228
|
-
maintainer's plea, a backdated git history built by a self-deleting kit), gate off.
|
|
229
|
-
**Zero frontier tampers in 20 dressed runs**, matching the undressed baselines, while
|
|
230
|
-
Haiku stayed a cheater at 8/10. The pre-registered manipulation check is the honest
|
|
231
|
-
part: only Opus actually read the dressing (8/10 runs, all ten honest) — so the Opus
|
|
232
|
-
null carries the weight and the Sonnet null is scoped as ambient-only. And the
|
|
233
|
-
sanctioned-edit variant (`harness/PREDICTION-self-authored.md`) — implement from an
|
|
234
|
-
ambiguous spec, then a prompt *licensing* correction of "intern-written" tests —
|
|
235
|
-
produced **zero test rewrites in 30 runs across three models**, refuting all three of
|
|
236
|
-
my rewrite bets, for a reason worth knowing: every model, implementing blind, chose
|
|
237
|
-
the same reading the tests encode. You cannot make an agent believe a test is wrong
|
|
238
|
-
by telling it tests might be wrong; the forced-conflict variant is the registered
|
|
239
|
-
follow-up.
|
|
240
|
-
|
|
241
|
-
**Can "closed vs open" be called in advance?** The two dimension studies
|
|
242
|
-
(`harness/PREDICTION-coverage.md`, `harness/PREDICTION-noverify.md`) tested whether a
|
|
243
|
-
rule's leak-prone axis can be identified a priori. Both a-priori calls **held
|
|
244
|
-
deterministically**: every enumerated hook-skip spelling denies while both
|
|
245
|
-
`core.hooksPath` forms allow (the axis called open is open — surveyed by six frontier
|
|
246
|
-
runs in reconnaissance, exploited by none), and all three coverage config-surface moves
|
|
247
|
-
block. The behavioral half delivered the study's best datum: told to get the coverage
|
|
248
|
-
gate green fast, **half of Haiku's runs opened with the per-path override** — the exact
|
|
249
|
-
config-surface move the registration predicted a token detector would miss — and the
|
|
250
|
-
shipped semantic detector caught every one. Zero tampers in 60 more guarded runs; one
|
|
251
|
-
false-positive interception recorded and kept (Fable extending a test-file import via
|
|
252
|
-
`sed`, denied by the shell-mutation heuristic mid-honest-fix — the FP record's first
|
|
253
|
-
live entry).
|
|
254
|
-
|
|
255
|
-
**And when the cheapest cheat is invisible to the gate?** The type-laundering study
|
|
256
|
-
(`harness/PREDICTION-asany.md`) probed the one seed whose cheapest cheat spellings are
|
|
257
|
-
deliberately uncovered — `: any` annotations are warn-only (`ts-any-launder`, the
|
|
258
|
-
FP-avoidance choice) and a bare assertion (`raw as Settings`) has no rule at all — with
|
|
259
|
-
a held-out runtime oracle hardened *before* any rep (without that hardening, nine
|
|
260
|
-
assertion tampers would have scored honest and the conclusion would have inverted).
|
|
261
|
-
Result: **Haiku laundered 10/10** — and the three runs denied on the guarded `as any`
|
|
262
|
-
cast **switched to the invisible assertion**, the first time a deny redirected an agent
|
|
263
|
-
to a better bypass instead of the fix. **Opus went 10/10 honest with zero gate
|
|
264
|
-
involvement** — the strongest disposition evidence measured — while **Fable tampered
|
|
265
|
-
6/10 and Sonnet, the default model, 10/10, both via the bare assertion and drawing
|
|
266
|
-
zero denies** (Sonnet's spelling every time: an inline structural assertion,
|
|
267
|
-
`as { retries: number; ... }` — the quickest silence that compiles carries no bypass
|
|
268
|
-
framing, and Sonnet's famous refusals turn out to be triggered by framing). The
|
|
269
|
-
completed frontier split — **Opus 0/10 · Fable 6/10 · Sonnet 10/10** — is the
|
|
270
|
-
published shape, with the standing self-measurement caveat on the Fable rows.
|
|
271
|
-
Twenty-six of forty runs tampered through the uncovered dimension: measured demand
|
|
272
|
-
that sent the `ts-any-launder` warn→block graduation question to a §7.A negatives
|
|
273
|
-
corpus — the honest fix itself narrows through `as Record<string, unknown>`, so a
|
|
274
|
-
naive assertion-blocker is an FP factory. **The corpus answered, and the question is
|
|
275
|
-
closed** (`harness/PREDICTION-launder-corpus.md`): on 1,227 real commits the
|
|
276
|
-
laundering spellings fire on legitimate work at 4–10× any deployable block threshold
|
|
277
|
-
(annotation-position `any`: 10.5% of commits; named-type assertions: 7.5%), so the
|
|
278
|
-
warn stays a warn permanently on syntactic evidence, no assertion rule gets built,
|
|
279
|
-
and the laundering class is owned by the layers where coverage is measured: the
|
|
280
|
-
held-out **oracle** (caught all 18 assertion tampers) and the **prompt** (the naming
|
|
281
|
-
sentence, 10/10 flip). For this class, detection is not the gate's layer —
|
|
282
|
-
verification and framing are, and the record says so with a denominator.
|
|
283
|
-
|
|
284
|
-
**Why not just a prompt line, or a 20-line grep hook?** Measured
|
|
285
|
-
(`harness/PREDICTION-baselines.md`), against a frozen naive hook and a one-sentence
|
|
286
|
-
prompt baseline — and two of my three bets lost, both against my own product's
|
|
287
|
-
framing. **One integrity sentence took ungated Haiku from 6/10 tampering to 1/10**:
|
|
288
|
-
the cheap prompt convention is a real mitigation and this README says so. The naive
|
|
289
|
-
hook converted denials nearly as well as the gate (4/5) on the crude-spelling seed —
|
|
290
|
-
but was tampered through twice: once by **hardcoding the implementation** (invisible
|
|
291
|
-
to any hook; the held-out oracle layer caught it), and once by a denied agent
|
|
292
|
-
switching to the node:test **options spelling** `{ skip: true }` — which the shipped
|
|
293
|
-
`test-skip` rule *also* missed until it shipped closed in 1.5.1. On 1,227 real merged
|
|
294
|
-
commits, the naive patterns would block **one in four legitimate commits (24.9%)**
|
|
295
|
-
versus the gate's 9.0% — a ~2.75× precision edge, not the ≥5× I bet, and that 9% on
|
|
296
|
-
unconfigured repos is published as the honest tuning-and-sign-off cost, not hidden.
|
|
297
|
-
The loop-layer deny alone was never the product; the measured margins are whole-class
|
|
298
|
-
spelling coverage, commit-layer precision plus sign-off machinery, and the layers no
|
|
299
|
-
hook has. Two follow-ups closed the loops this study opened. The sentence was then
|
|
300
|
-
aimed at the regime where the gate is blind — Sonnet's 10/10 assertion-laundering
|
|
301
|
-
seed — and produced a **complete flip: 10/10 honest** (`harness/PREDICTION-baselines.md`,
|
|
302
|
-
addendum): naming the act is the intervention Sonnet's framing-triggered honesty
|
|
303
|
-
runs on. And the promised **tuned-gate numbers** now exist
|
|
304
|
-
(`harness/PREDICTION-tuned-fp.md`): 9.0% unconfigured → **6.0%** with a one-line
|
|
305
|
-
severity tune → roughly one sign-off decision per 17 commits residual. The same study
|
|
306
|
-
shipped per-rule `exclude` globs (1.6.0, with the policy-weakening symmetry built in)
|
|
307
|
-
and then measured them honestly: **zero effect on this corpus** — these repos' casts
|
|
308
|
-
live in source, not tests, so the zod-recommended path scoping removed nothing here;
|
|
309
|
-
the null is published next to the feature.
|
|
310
|
-
|
|
311
|
-
**What about just making the tests read-only?** The most common objection got its
|
|
312
|
-
pre-registered measurement (`harness/PREDICTION-capability.md`) — starting with the
|
|
313
|
-
mechanics the folk advice misses: `chmod` does not bind a root agent and never
|
|
314
|
-
prevented `rm` anyway (deletion is a *directory*-write right); the workable barrier
|
|
315
|
-
is the immutable bit. Result: **hard immutability governed the crude-tamper seed
|
|
316
|
-
completely — 10/10 honest, four EPERM hits all converting to the real fix** (a bare
|
|
317
|
-
permission error transferred the deny→fix dynamic with no remediation text). Then
|
|
318
|
-
the twist: adding a logged, auto-granted **elevation script made it worse (8/10)** —
|
|
319
|
-
both extra tampers walked through the front door with written confessions
|
|
320
|
-
("*Remove failing test file to unblock build*"), the sanctioned-path thesis firing
|
|
321
|
-
a fourth time. The lock is easy; the grantor is the hard part — which is the case
|
|
322
|
-
for elevation flows that end at a human, and exactly what the gate's sign-off label
|
|
323
|
-
already is. That seed now carries all six governance configurations on identical
|
|
324
|
-
pre-registered ground: nothing 6/10 tampers, one sentence 1/10, naive hook 2/10,
|
|
325
|
-
the gate 0/10, immutable 0/10, immutable+auto-grant 2/10 — with the standing scope
|
|
326
|
-
that none of the capability rows touch the semantic classes, by construction.
|
|
327
|
-
Running ledger: **70 pre-registered bets — 29 refuted, 37 held, 4 open or
|
|
328
|
-
unresolvable** — taskbench's B2 (prevention on real repositories) is the newest
|
|
329
|
-
refutation, published beside the bet it refutes.
|
|
330
|
-
|
|
331
|
-
**What about false positives?** `snapshot-rewrite` was swept over **1,652 real mainline
|
|
332
|
-
commits** (prettier, jest, docusaurus, immer): 216 touched snapshots, all legitimately —
|
|
333
|
-
which is exactly why that rule is a `warn` asking for human confirmation, not a block,
|
|
334
|
-
and why the study is committed (`harness/fp-study/`) rather than summarized away. The
|
|
335
|
-
study's narrow signal — snapshot-only diffs, ~0.06% of audited mainline commits but
|
|
336
|
-
7/7 of observed tampers — now ships as the distinct `snapshot-only-rewrite` warn, the
|
|
337
|
-
graduation candidate.
|
|
158
|
+
## Four layers, different jobs
|
|
159
|
+
|
|
160
|
+
| layer | the question it answers |
|
|
161
|
+
| --- | --- |
|
|
162
|
+
| **Change semantics** | Does a proposed or recorded change match a known weakening class? (PreToolUse deny, staged diff, range diff) |
|
|
163
|
+
| **Effect** | What did the protected state actually become, however the mutation arrived? (drift check, fs-event observer, Stop sweep) |
|
|
164
|
+
| **Outcome** | Does the candidate still pass when protected verification state from the base is restored? (pristine verification — no diff interpreted at all) |
|
|
165
|
+
| **Authority** | Where is the final verdict rendered, and what assumptions make that verdict authoritative? (run envelope, CI, SHA-bound human sign-off) |
|
|
166
|
+
|
|
167
|
+
Sixteen rules are specified and fourteen ship (see the table in
|
|
168
|
+
[SPEC.md](./SPEC.md)). The families: test protection (`test-deletion`,
|
|
169
|
+
`test-skip`, `test-content-removal`), verification-signal protection
|
|
170
|
+
(`coverage-lowering`, `snapshot-rewrite`, `snapshot-only-rewrite`), suppression
|
|
171
|
+
(`ts-any-cast`, `ts-any-launder`, `lint-suppression`), pipeline protection
|
|
172
|
+
(`ci-tampering`, `hook-tampering`, `no-verify`), and the effect/outcome layers
|
|
173
|
+
(`transient-protected-mutation`, `pristine-verification`, plus the `run` envelope).
|
|
174
|
+
Some ambiguous syntactic classes deliberately remain warnings or unimplemented —
|
|
175
|
+
`assertion-weakening` and `guard-removal` are reserved names with no detector, and
|
|
176
|
+
`ts-any-launder` is a permanent warn — rather than being promoted to blockers
|
|
177
|
+
without precision evidence.
|
|
178
|
+
|
|
179
|
+
## What this does not establish
|
|
180
|
+
|
|
181
|
+
- **Not a correctness oracle.** Pristine verification can only re-run tests that
|
|
182
|
+
exist in the tree. Round 2's designed-in blind spot: on tasks with withheld
|
|
183
|
+
test cases, an agent that honestly half-fixes the visible cases gets a genuine
|
|
184
|
+
visible green and is caught only by the withheld half — no tampering involved.
|
|
185
|
+
A hardcoded implementation is likewise invisible to any diff rule.
|
|
186
|
+
- **Not an OS sandbox, a network firewall, a secret manager, or a replacement for
|
|
187
|
+
branch protection.** It enforces verification integrity; it does not confine
|
|
188
|
+
the process.
|
|
189
|
+
- **The intended deployment is defence in depth:** instructions naming the
|
|
190
|
+
forbidden move + capability restriction + Tamperward + independent held-out
|
|
191
|
+
tests + authoritative CI. Each layer covers classes the others measurably miss.
|
|
192
|
+
|
|
193
|
+
**A word on the words:** *tamper* and *honest* are operational labels on
|
|
194
|
+
artifacts — a protected asset changed, or an independent oracle failed, while the
|
|
195
|
+
visible check went green — decided deterministically. They are not claims about
|
|
196
|
+
any model's intent.
|
|
197
|
+
|
|
198
|
+
## Research discipline
|
|
199
|
+
|
|
200
|
+
Predictions are registered before counted runs, with numeric bets and explicit
|
|
201
|
+
losing conditions, and the git history proves the order. Seeds, pools, frozen
|
|
202
|
+
analysis scripts, transcripts, and deviation ledgers are committed. A pool used
|
|
203
|
+
to develop detectors is spent — validation requires a fresh draw (round 1's
|
|
204
|
+
repositories are development data now; round 2 drew new ones). Run counts are
|
|
205
|
+
reported with their honest evidence unit: repeated runs on one seed measure that
|
|
206
|
+
seed's stochasticity, and the independent evidence behind the synthetic studies
|
|
207
|
+
is the count of distinct seed configurations, not the run count. Losing
|
|
208
|
+
predictions and corrections remain in the public record rather than being
|
|
209
|
+
removed after the result is known; the series and its errata carry the ledger
|
|
210
|
+
and its totals.
|
|
211
|
+
|
|
212
|
+
**[The research series](./docs/blog/index.md)** ·
|
|
213
|
+
**[The harness](./harness/)** ·
|
|
214
|
+
**[Errata](./docs/blog/errata.md)**
|
|
215
|
+
|
|
216
|
+
## The architecture was earned by failures
|
|
217
|
+
|
|
218
|
+
| study / version | consequence |
|
|
219
|
+
| --- | --- |
|
|
220
|
+
| Early seed studies | The proposal gate: deny the known weakening classes before the tool call runs, with the deny message as the correction signal |
|
|
221
|
+
| Blind-dimension studies | Coverage ≠ completeness: when the cheapest cheat sat outside detector coverage, agents took it — so an orthogonal verification layer matters more than more rules |
|
|
222
|
+
| Round 1, p = 1.0 | The 53 real trajectories became the evasion corpus — evasion classes the detectors never imagined, extracted with full forensics |
|
|
223
|
+
| v1.7.0 | `test-content-removal`: the row-deletion / assertion-stripping / setup-gutting family that carried round 1's gated false greens, corpus-priced on 2,304 real commits before shipping as block |
|
|
224
|
+
| v1.8.0 | The effect layer: per-call protected-tree drift check plus the `tamperward watch` fs-event observer — judge what the files became, not the command spelling |
|
|
225
|
+
| v1.9.0 | `tamperward verify`: the round-1 pristine oracle productized — visible-green + pristine-red is a masked failure, and "could not verify" fails closed |
|
|
226
|
+
| Round 2, 9–0 | On 22 fresh repositories, v1.9.0 produced b=9/c=0 false-green discordance (p = 0.0039); four gated false greens remained, through two documented residual mechanisms |
|
|
227
|
+
| cssstyle lifecycle escape (round 2) | The `tamperward run` envelope: the gate denied 42 mutations and verify refused all 25 stop attempts, and the runtime completed the session anyway over a masked tree — hooks decide, runtimes terminate, so the adjudication boundary moved outside the runtime |
|
|
228
|
+
| v1.14.1, article audit | `tamperward verify` kept agent-added protected files in the pristine run on the premise that they "only add strictness". An added `conftest.py` could deselect the restored base tests by node id, so a masked failure reported VERIFIED and the envelope printed GREEN MEANS GREEN over an unfixed bug. Added protected files are now removed; PoC and mutation-checked regression committed |
|
|
229
|
+
| v1.10.1–v1.14.0, owner + two-pass external review | Frozen entry-time policy and verifier, entry-SHA ancestry enforcement, quiescence guard, dependency-tree fingerprint, a CI verify step in the generated workflow, the gate pinned to its own version in CI, and SHA-bound sign-off labels |
|
|
230
|
+
|
|
231
|
+
Each row's primary artifact: [CHANGELOG.md](./CHANGELOG.md), [SPEC.md](./SPEC.md),
|
|
232
|
+
and the posts in [docs/blog/](./docs/blog/index.md).
|
|
338
233
|
|
|
339
234
|
## Stability
|
|
340
235
|
|
|
@@ -352,7 +247,7 @@ via npm trusted publishing with SLSA provenance. Full rule:
|
|
|
352
247
|
|
|
353
248
|
```bash
|
|
354
249
|
npm install && npm run build # bundles the CLI to dist/cli/index.js
|
|
355
|
-
npm test #
|
|
250
|
+
npm test # 365 tests at v1.14.1 — parser, detectors, engine, policy, renderers
|
|
356
251
|
npm run typecheck
|
|
357
252
|
```
|
|
358
253
|
|
package/dist/cli/index.js
CHANGED
|
@@ -2900,6 +2900,32 @@ import { chmodSync as chmodSync2, cpSync, lstatSync, mkdirSync as mkdirSync6, mk
|
|
|
2900
2900
|
import { tmpdir as tmpdir2 } from "node:os";
|
|
2901
2901
|
import { dirname as dirname4, join as join9, resolve } from "node:path";
|
|
2902
2902
|
var OVERLAY_CLASSES = ["tests", "snapshots", "config"];
|
|
2903
|
+
var VERIFICATION_SURFACE = [
|
|
2904
|
+
// Python / pytest — read from the rootdir, and conftest at any depth
|
|
2905
|
+
"**/conftest.py",
|
|
2906
|
+
"**/pytest.ini",
|
|
2907
|
+
"**/.pytest.ini",
|
|
2908
|
+
"**/setup.cfg",
|
|
2909
|
+
"**/tox.ini",
|
|
2910
|
+
"**/pyproject.toml",
|
|
2911
|
+
// JavaScript / TypeScript runners and the transforms they load
|
|
2912
|
+
"**/jest.config.*",
|
|
2913
|
+
"**/jest.setup.*",
|
|
2914
|
+
"**/vitest.config.*",
|
|
2915
|
+
"**/vitest.workspace.*",
|
|
2916
|
+
"**/.mocharc.*",
|
|
2917
|
+
"**/karma.conf.*",
|
|
2918
|
+
"**/babel.config.*",
|
|
2919
|
+
"**/.babelrc*",
|
|
2920
|
+
"**/vite.config.*",
|
|
2921
|
+
"**/.swcrc",
|
|
2922
|
+
"**/package.json",
|
|
2923
|
+
// Ruby / PHP / .NET
|
|
2924
|
+
"**/.rspec",
|
|
2925
|
+
"**/phpunit.xml",
|
|
2926
|
+
"**/phpunit.xml.dist",
|
|
2927
|
+
"**/*.runsettings"
|
|
2928
|
+
];
|
|
2903
2929
|
function git2(args, cwd) {
|
|
2904
2930
|
return execFileSync4("git", args, { cwd, encoding: "utf8", maxBuffer: 1 << 28 });
|
|
2905
2931
|
}
|
|
@@ -2948,7 +2974,7 @@ function dropSymlink(p) {
|
|
|
2948
2974
|
}
|
|
2949
2975
|
function overlayPristine(cwd, base, dest, policy) {
|
|
2950
2976
|
const atBase = git2(["ls-tree", "-r", "--name-only", "-z", base], cwd).split("\0").filter(Boolean);
|
|
2951
|
-
const isOverlay = (p) => OVERLAY_CLASSES.some((c) => isProtected(p, policy, c));
|
|
2977
|
+
const isOverlay = (p) => OVERLAY_CLASSES.some((c) => isProtected(p, policy, c)) || matchesAny(p, VERIFICATION_SURFACE);
|
|
2952
2978
|
let restored = 0;
|
|
2953
2979
|
const baseProtected = /* @__PURE__ */ new Set();
|
|
2954
2980
|
for (const rel of atBase) {
|
|
@@ -2961,10 +2987,14 @@ function overlayPristine(cwd, base, dest, policy) {
|
|
|
2961
2987
|
writeFileSync6(out2, content);
|
|
2962
2988
|
restored++;
|
|
2963
2989
|
}
|
|
2964
|
-
let
|
|
2990
|
+
let removed = 0;
|
|
2965
2991
|
const inCopy = git2(["ls-files", "-z", "--cached", "--others", "--exclude-standard"], cwd).split("\0").filter(Boolean);
|
|
2966
|
-
for (const rel of inCopy)
|
|
2967
|
-
|
|
2992
|
+
for (const rel of inCopy) {
|
|
2993
|
+
if (!isOverlay(rel) || baseProtected.has(rel)) continue;
|
|
2994
|
+
rmSync2(join9(dest, rel), { force: true });
|
|
2995
|
+
removed++;
|
|
2996
|
+
}
|
|
2997
|
+
return [restored, removed];
|
|
2968
2998
|
}
|
|
2969
2999
|
function runSuite(dir, cmd, budgetSecs) {
|
|
2970
3000
|
const t0 = Date.now();
|
|
@@ -3015,13 +3045,13 @@ function runVerify(opts) {
|
|
|
3015
3045
|
const visDir = join9(work, "visible");
|
|
3016
3046
|
const priDir = join9(work, "pristine");
|
|
3017
3047
|
let restored = 0;
|
|
3018
|
-
let
|
|
3048
|
+
let removedAdded = 0;
|
|
3019
3049
|
try {
|
|
3020
3050
|
mkdirSync6(visDir);
|
|
3021
3051
|
mkdirSync6(priDir);
|
|
3022
3052
|
materialize(cwd, visDir);
|
|
3023
3053
|
materialize(cwd, priDir);
|
|
3024
|
-
[restored,
|
|
3054
|
+
[restored, removedAdded] = overlayPristine(cwd, base, priDir, policy);
|
|
3025
3055
|
} catch (e) {
|
|
3026
3056
|
if (!opts.keep) rmSync2(work, { recursive: true, force: true });
|
|
3027
3057
|
out2(`verify: could not materialize (${e instanceof Error ? e.message : String(e)}) \u2014 failing closed`);
|
|
@@ -3055,7 +3085,7 @@ function runVerify(opts) {
|
|
|
3055
3085
|
visible: { exit: visible.exit, secs: visible.secs },
|
|
3056
3086
|
pristine: { exit: pristine.exit, secs: pristine.secs },
|
|
3057
3087
|
protected_restored: restored,
|
|
3058
|
-
|
|
3088
|
+
added_protected_removed: removedAdded,
|
|
3059
3089
|
...opts.keep ? { workdir: work } : {}
|
|
3060
3090
|
})
|
|
3061
3091
|
);
|
|
@@ -3067,7 +3097,10 @@ function runVerify(opts) {
|
|
|
3067
3097
|
BUDGET_EXCEEDED: `budget exceeded (${budget}s): could not verify \u2014 failing closed, not open.`
|
|
3068
3098
|
};
|
|
3069
3099
|
out2(`tamperward verify \u2014 ${lines[verdict2]}`);
|
|
3070
|
-
if (
|
|
3100
|
+
if (removedAdded > 0)
|
|
3101
|
+
out2(
|
|
3102
|
+
`(${removedAdded} protected file(s) added since ${base.slice(0, 10)} were removed from the pristine run: the pristine tree carries exactly the base's protected surface.)`
|
|
3103
|
+
);
|
|
3071
3104
|
}
|
|
3072
3105
|
return code;
|
|
3073
3106
|
}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "tamperward",
|
|
3
|
-
"version": "1.14.
|
|
3
|
+
"version": "1.14.2",
|
|
4
4
|
"description": "The deterministic agent-integrity gate. One ruleset, evaluated on the actual diff/commands as a verdict, enforced everywhere a change can be made.",
|
|
5
5
|
"license": "Apache-2.0",
|
|
6
6
|
"author": "hexrift",
|