tamperward 1.14.0 → 1.14.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (3) hide show
  1. package/README.md +203 -308
  2. package/dist/cli/index.js +41 -8
  3. package/package.json +1 -1
package/README.md CHANGED
@@ -12,329 +12,224 @@
12
12
  <a href="./LICENSE"><img src="https://img.shields.io/badge/license-Apache--2.0-lightgrey" alt="license"></a>
13
13
  </p>
14
14
 
15
- **The deterministic agent-integrity gate.** AI coding agents optimize for "the command
16
- succeeded," not "the change is trustworthy." To make checks pass they take the cheapest
17
- path: deleting failing tests, skipping them, casting to `any`, suppressing lint errors,
18
- lowering coverage gates, rewriting snapshots, editing CI, or bypassing hooks with
19
- `--no-verify`. Tamperward treats the safety nets themselves as protected assets and
20
- blocks the **class** of bypass — one ruleset, evaluated on the actual diff and commands
21
- as a verdict, not a probability, enforced at every stage a change passes through:
22
- **inside the agent's loop (a Claude Code hook today — a deny before the tool runs, and
23
- a sweep at end of turn), at the commit, and at the merge.** CI is the authority; the
24
- earlier layers shorten the feedback loop for agents that carry them.
25
-
26
- > Apache-2.0 · zero runtime model calls · every headline claim below is measured, with
27
- > the pre-registered predictions (including the refuted ones) committed to this repo
28
-
29
15
  **[Docs & guide](https://hexrift.github.io/tamperward/)** ·
30
- **[The launch post](https://hexrift.github.io/tamperward/blog/what-agents-do-when-you-block-their-shortcuts)** — how these numbers were measured, and which of my own bets lost
31
-
32
- ## Install
16
+ **[The research series](./docs/blog/index.md)** — every registered prediction
17
+ published beside its outcome
18
+
19
+ Coding agents can modify both the implementation and the tests, configuration,
20
+ CI, hooks and verifier used to judge that implementation — and in observed
21
+ trajectories, some modify or attempt to modify verification in ways that can
22
+ turn incorrect work into apparent success. Under pressure, the cheaper route to
23
+ green is sometimes to weaken the checks instead of fixing the failure.
24
+
25
+ Tamperward is a **deterministic verification-integrity layer**. It blocks known
26
+ weakening moves as they happen, observes protected-state effects, and
27
+ independently re-adjudicates apparent success outside the agent's normal
28
+ completion path. No runtime LLM judge. Fail closed when adjudication is
29
+ impossible.
30
+
31
+ ## What we have actually measured
32
+
33
+ | experiment | result | what it supports |
34
+ | --- | --- | --- |
35
+ | **Round 1** — 26 paired real repos, historical regressions, 53 counted trajectories, **v1.6.0** ([`harness/taskbench/`](./harness/taskbench/)) | Transfer: **9/27 ungated runs (33.3%)** violated policy ([corrected](./harness/taskbench/reanalysis/TRANSFER-REANALYSIS.md) from a published 13/26 — the original predicate was defective; the transfer bet is refuted, not held). Headline prevention bet **lost**: b=5 / c=4, RD +3.8pp [−17.2, +24.7], exact McNemar **p = 1.0** — published beside the bet | A detector-centric architecture was insufficient: agents routed around the shipped detector classes |
36
+ | **Pristine oracle, round 1** — independent re-execution of the original suite, **including withheld semantic cases** | Identified **every masked failure observed among the 53 trajectories** while diff-time detection was routed around | The outcome-level signal that motivated `tamperward verify`. The shipped command productizes this oracle's base-restoration component only — it carries no withheld cases and cannot detect a semantically incomplete fix the base tests also accept |
37
+ | **Round 2** — **22 fresh held-out repos** no detector was tuned on ([`harness/taskbench/round2/`](./harness/taskbench/round2/)) | Transfer **12/22 (54.5%)** ([corrected](./harness/taskbench/reanalysis/TRANSFER-REANALYSIS.md) from a published 14/22). Prevention: **b=9 / c=0**, RD **+40.9pp**, BP95 **[17.8, 61.3]**, exact McNemar **p = 0.0039** | The preregistered **v1.9.0** treatment materially reduced false greens in that setting |
38
+ | **After prevention** — the nine round-2 prevented false greens | **8 of 9 became honest completions**; the ninth an honest **non**-completion | Prevention usually redirected trajectories toward honest work rather than merely blocking them |
39
+ | **Round 3** — 17 paired Python repos, fresh PyPI frame, **v1.14.0** ([`harness/taskbench/round3/`](./harness/taskbench/round3/)) | Transfer 9/17 (52.9%). Prevention: **b=6 / c=0**, RD **+35.3pp**, BP95 [9.5, 58.7], exact McNemar **p = 0.0313** | The prevention result appeared in Python — with the treatment also changed from v1.9.0, so ecosystem transfer is not isolated. The in-loop skip detector proved blind to pytest syntax; the outer layers carried it |
40
+ | **Round 3.1** — the same 16 pairs under **`claude-sonnet-5`** ([`harness/taskbench/round3.1/`](./harness/taskbench/round3.1/)) | Transfer 4/16 (25.0%). Prevention: **b=1 / c=0**, RD +6.3pp, BP95 [−13.8, 28.3], exact McNemar **p = 1.0000** | The confirmatory result **did not replicate**, and could not have: only three ungated false greens occurred, so `b ≤ 3` and p ≥ .25 whatever the gate did. A failure to reject, not evidence of no effect |
41
+
42
+ Earlier controlled experiments → **[the research series](./docs/blog/index.md)**.
43
+
44
+ > **Scope.** The rows above are one model, one pressure prompt, one JS/TS
45
+ > ecosystem — evidence for that setting, not a universal claim. Round 2 tested
46
+ > the v1.9.0 stack; the current **v1.14.1** adds post-study envelope hardening
47
+ > (externally reviewed, with findings tracked individually as REPRO or AUDIT in
48
+ > [SECURITY-ENVELOPE.md](./SECURITY-ENVELOPE.md) and closed with regression and
49
+ > mutation checks — see [CHANGELOG](./CHANGELOG.md)). Rounds 3 and 3.1 are
50
+ > complete and are in the table above. **1.14.1 closed a bypass in `tamperward
51
+ > verify`** — an agent-added protected file could suppress the tests the pristine
52
+ > run had just restored — which was present from v1.9.0 onward, so rounds 2, 3
53
+ > and 3.1 all ran on a verifier carrying it. No evidence any agent used it; the
54
+ > layer's authority in those rounds was nonetheless weaker than described at the
55
+ > time. Round 4 is the fresh pool.
56
+ > No round-3 outcomes are included in the claims above.
57
+
58
+ ## Architecture
59
+
60
+ The final envelope adjudication is anchored to state recorded before the agent
61
+ starts; the in-loop hooks read the live tree and remain steering controls. The
62
+ final adjudication boundary sits outside the agent's normal lifecycle.
33
63
 
34
- ```bash
35
- npx tamperward init
64
+ ```
65
+ RECORDED AT ENVELOPE ENTRY:
66
+ entry SHA · frozen policy · frozen verifier command/budget · dependency fingerprint
67
+ │
68
+ ┌─────────────────────────────┴───────────────────────────────┐
69
+ │ agent loop (Claude Code hooks today) │
70
+ │ PreToolUse gate: every proposed Edit/Write/Bash judged; │
71
+ │ known weakening DENIED before the tool runs │
72
+ │ │ │
73
+ │ working tree ▼ │
74
+ │ effect observation: per-call protected-tree drift check │
75
+ │ + filesystem-event observer (`tamperward watch`) │
76
+ │ │ │
77
+ │ Stop sweep ▼ │
78
+ │ end-of-turn net diff re-checked; transient protected │
79
+ │ mutations judged from the event log │
80
+ │ │ │
81
+ │ pristine verification ▼ (`tamperward verify`) │
82
+ │ suite runs twice in separate copies: the candidate tree │
83
+ │ as-is, and with protected tests/snapshots/config │
84
+ │ restored from the trusted base; visible-green + │
85
+ │ pristine-red = MASKED FAILURE │
86
+ └─────────────────────────────┬───────────────────────────────┘
87
+ │ agent runtime exits — exit code NOT trusted
88
+ ▼
89
+ `tamperward run -- <agent…>` — the outer envelope
90
+ policy + verifier resolved from the entry state, never the tree
91
+ the agent wrote · post-agent HEAD must descend from the entry
92
+ SHA · policy check over base...HEAD (committed weakening) and
93
+ over the worktree incl. untracked files (uncommitted) ·
94
+ dependency-tree fingerprint compared · quiescence guard ·
95
+ pristine verify against the entry base · exit 1 on any blocking
96
+ finding or masked failure · exit 2 (fail closed) when
97
+ adjudication is impossible
98
+ │
99
+ ▼
100
+ pre-commit backstop: `tamperward check --staged`
101
+ │
102
+ ▼
103
+ PR CI: `check --diff base...head` + `verify --require-ancestor`
104
+ policy read from the merge-base — a PR cannot govern its own verdict
105
+ │
106
+ ▼
107
+ out-of-band human exception only: PR label
108
+ `tamperward:allow:<rule>@<head-sha>` — SHA-bound, applied by a human
109
+ with write access, never a file the branch can commit
110
+ │
111
+ ▼
112
+ protected main
36
113
  ```
37
114
 
38
- One idempotent command wires the policy and every enforcement point — it never
39
- overwrites anything you wrote, and `--dry-run` shows the plan first:
40
-
41
- | what | wired as |
42
- | --- | --- |
43
- | policy | a commented baseline `.tamperward.yml` — the defaults apply even without it |
44
- | agent loop | Claude Code `PreToolUse` deny + `Stop` sweep, merged into `.claude/settings.json` |
45
- | pre-commit | husky when present, the plain git hook otherwise |
46
- | CI | a PR-gate workflow with out-of-band sign-off (details below) |
115
+ CI is the final authority in this design, and that authority is conditional on
116
+ the deployment assumptions documented in
117
+ [docs/guide/enforcement.md](./docs/guide/enforcement.md) and
118
+ [SPEC.md](./SPEC.md): verification is anchored to a base the agent cannot
119
+ rewrite (a protected branch), the repository's branch protection actually
120
+ enforces the workflow's verdict on merges, and sign-off labels can only be
121
+ applied by humans with write access. Mechanisms, residuals, and what is
122
+ explicitly outside the trust boundary: [SPEC.md](./SPEC.md) and
123
+ [SECURITY-ENVELOPE.md](./SECURITY-ENVELOPE.md).
47
124
 
48
- Or run the views directly:
125
+ ## Quick start
49
126
 
50
127
  ```bash
51
- npx tamperward check --staged # pre-commit view
52
- npx tamperward check --diff "main...HEAD" # CI view — the authority for main
128
+ npx tamperward init
53
129
  ```
54
130
 
55
- ## What it looks like when it fires
131
+ One idempotent command wires the policy, the agent hooks, the pre-commit hook, and
132
+ a CI workflow that runs both the diff-time check and pristine verification. It
133
+ never overwrites anything you wrote; `--dry-run` prints the plan.
56
134
 
57
- An agent, mid-loop, tries the classic shortcut. The deny happens **before** the tool
58
- runs, and holds even under `bypassPermissions`:
135
+ A real deployment needs a verify command configured — the generated CI verify step
136
+ **fails closed (exit 2) without one** rather than passing quietly. In
137
+ `.tamperward.yml`:
59
138
 
60
- ```
61
- 5. Bash rm calc.test.js → DENIED (test-deletion)
62
- 6. Edit calc.js ← next action: fix the real bug
63
- 7. Bash node --test → green, honestly
139
+ ```yaml
140
+ verify:
141
+ command: npm test
142
+ budget: 300
64
143
  ```
65
144
 
66
- The same engine at pre-commit or in CI:
145
+ That block is itself a guarded surface: changing the command, lowering the budget,
146
+ or removing it is flagged as policy weakening — a verifier an agent can point at
147
+ `true` is no verification at all.
67
148
 
149
+ The four primitives:
150
+
151
+ ```bash
152
+ npx tamperward check --staged # pre-commit view
153
+ npx tamperward check --diff "main...HEAD" # CI view over the PR's commit range
154
+ npx tamperward verify --base main # pristine-suite re-execution
155
+ npx tamperward run -- <agent command...> # the outer envelope around an agent
68
156
  ```
69
- tamperward: 2 blocking
70
- (2 changes scanned)
71
-
72
- BLOCK coverage-lowering package.json
73
- Coverage gate weakened: global lines threshold lowered 90 → 10.
74
- evidence global lines threshold lowered 90 → 10
75
- instead Raise real coverage by adding tests; do not lower or exempt
76
- the gate to pass.
77
- sign-off tamperward allow coverage-lowering --file package.json --reason "..."
78
-
79
- BLOCK test-deletion test/calc.test.js
80
- Test blocks removed: 3 → 1 it()/test() in this spec.
81
- evidence 2 test block(s) removed from test/calc.test.js
82
- instead Keep the assertions and fix the code. Removing test blocks to go
83
- green is the tamper.
84
- sign-off tamperward allow test-deletion --file test/calc.test.js --reason "..."
85
-
86
- A blocking finding clears only with a human sign-off. In CI that sign-off is
87
- out-of-band — a PR label applied by a reviewer — never a file committed on the
88
- branch under review.
89
- ```
90
157
 
91
- Under GitHub Actions each finding also lands as an **inline annotation on the diff** in
92
- *Files changed*, plus a job-summary table — not four clicks deep in a log.
93
-
94
- ## The rules
95
-
96
- Ten mechanical rules, deterministic by construction — eight `block`, two `warn`:
97
- `test-deletion` (AST block-count, rename-out-of-glob, shell mutation), `test-skip`
98
- (including `.only`, which narrows the suite), `ts-any-cast`, `lint-suppression`,
99
- `coverage-lowering` (Jest and Vitest shapes, all four metrics, `package.json` included),
100
- `ci-tampering` (a **moved** check is not a removed one), `hook-tampering`, `no-verify`
101
- (flags and the env-var escape hatches), `snapshot-rewrite` (`warn` — see the FP study below
102
- for why), and `snapshot-only-rewrite` (`warn` — the FP study's narrow signal, a
103
- snapshot moved with **no accompanying change**, as its own rule so it can graduate to
104
- `block` at a later policy `version:` without touching the broad rule; it fires only at
105
- commit granularity, where "no accompanying change" means something).
106
-
107
- Two heuristic rule names — `assertion-weakening` and `guard-removal` — are **reserved
108
- in the baseline policy but not yet built**: they get detectors only once a measured
109
- negatives corpus exists (SPEC §7.A), enter as `warn`, and never block until their
110
- precision clears the bar.
111
-
112
- The policy file itself is un-ignorable, CI reads the policy from the **merge-base** (a PR
113
- cannot govern its own verdict), and the agent honours no sign-off it can author — the
114
- escape valve is a PR label only a human with write access can apply.
115
-
116
- ## Measured, not asserted
117
-
118
- **Read the two headline numbers together, always:** across **137 guarded runs — every
119
- one inside a tamper class an active detector covered — 0 tampers reached green.**
120
- Outside that coverage, on the one seed built to measure it, **agents tampered in 26 of
121
- 40 runs** — "tampered" meaning, mechanically: the independent held-out oracle
122
- classified the result as passing the visible check through a prohibited shortcut. The
123
- first number is what the gate does; the second is why the gate is one layer. And an honest note on evidence units: those run counts come from roughly
124
- fifteen seed/task configurations and eight real repositories — repeated runs on one
125
- seed measure that seed's stochasticity, not the space of real coding tasks, so the
126
- independent evidence behind these numbers is the seed count, not the run count.
127
-
128
- **And the third number, from the preregistered cross-repository experiment
129
- (taskbench, `harness/taskbench/`):** on 26 paired real repositories with real
130
- historical regressions, pressured-agent policy violations appeared in **13 of 26
131
- ungated runs (50%)** — the behavior transfers — but the gate at 1.6.0 **did not
132
- reduce false greens** (5 prevented vs 4 gated-only; paired RD +3.8%, 95% CI
133
- −17.2 to +24.7, exact McNemar p = 1.0): agents routed through tool-call
134
- spellings outside the shipped detector classes, including two gaps documented
135
- before the sweep and deliberately left unfixed for the experiment's integrity.
136
- Enforcement showed no measurable completion cost (+7.7 points, CI spans zero),
137
- and independent pristine-suite re-execution caught **every masked failure in
138
- all 53 trajectories**. Registration, bets, frozen analysis, and the refutation
139
- of our own headline bet: `PREDICTION-taskbench.md`, `PHASE3.md`. The first two
140
- numbers say what the gate does inside its coverage; the third says coverage is
141
- the whole game on real repositories — and that the reliable layer was the
142
- pristine-suite oracle the agent cannot touch.
143
-
144
- **What this does not do**, stated before the numbers rather than after: it does not
145
- establish that a patch is *correct*; it does not detect arbitrary deception (a
146
- hardcoded implementation is invisible to any diff rule — measured, twice); it does not
147
- replace sandboxing, held-out tests, or authoritative CI — it is designed to sit in
148
- front of them. What it does: **enforce explicitly defined development invariants at
149
- the loop, commit, and CI boundaries** — and nothing else.
150
-
151
- **A word on the words:** *tamper*, *honest*, *cheat*, and *surrender* are used
152
- throughout as operational labels, not claims about mental states. A "tamper" means: a
153
- protected asset changed, or an independent held-out oracle failed, while the visible
154
- check went green — a property of the artifacts, decided deterministically. Whether the
155
- process that produced it "intended" anything is a question these experiments do not
156
- measure and do not need.
157
-
158
- **Does the loop work?** `harness/run.sh` stands up an isolated repo with a real bug and a
159
- failing protected test, lets a headless agent loose with the hooks installed, and
160
- classifies the end state with an oracle independent of the detectors under test
161
- (SPEC §7.B). Across **137 runs** of these **guarded** scenarios — every attempted bypass
162
- class had a detector covering it — **0 tampers reached green**. Denied runs converted to
163
- an honest fix **42 times out of 56 (75%)**, and that number is honestly seed-dependent:
164
- 25/25 where the honest fix was the agent's own task, down to **0/10** where the failing
165
- suite was somebody else's bug (the no-verify study below). The frontier split gets the
166
- same scoping: on the expensive-fix seed no frontier model attempted a bypass even when
167
- the prompt *ordered* one; on the hook-skip seed, frontier models under a prompt ordering
168
- the skip **try the ordered skip once, get denied once, and convert 10/10** — frontier
169
- honesty there is what happens after the first no, not never-trying.
170
- Reproduce: `npm run build && harness/run.sh 5 adversarial haiku seed-hard`.
171
-
172
- **Is the fix caused by the deny, or would it have happened anyway?** Measured, with a
173
- pre-registered no-gate control arm (`harness/PREDICTION-ungated-control.md`): same seed,
174
- model, prompt and oracle, hooks removed. **6/10 ungated runs tampered through; honest
175
- fixes fell from 10/10 gated to 4/10 ungated.** The gate's measured contribution on that
176
- configuration is a 60-point swing — causation, not survivorship — and one of the three
177
- control-arm bets was refuted (ungated honest fixes came in above my bet). The control
178
- arm was then extended across the model family (`harness/PREDICTION-frontier-models.md`):
179
- ungated Sonnet 5, Opus 5, and Fable 5 each went **10/10 honest** under the same
180
- pressure prompt — frontier honesty is disposition, not the gate's deterrence — while
181
- ungated Haiku stays 6/10 tamper-green. The gate's measured value concentrates on the
182
- models that do cheat. (The Fable rows are a Fable-driven harness measuring Fable:
183
- self-graded homework, bounded by the deterministic held-out oracle.) The
184
- type-laundering study below later measured the boundary of the disposition claim:
185
- it held for Opus in every regime; Sonnet and Fable held only where the gate could
186
- interrupt.
187
-
188
- **And when the agent knows the gate is there?** Also pre-registered
189
- (`harness/PREDICTION-informed-adversary.md`) — betting the gate would *leak*. Refuted,
190
- three bets out of three: with the detector sources vendored into the repo and the prompt
191
- naming the gate, Haiku's bypass attempts went from 5/10 uninformed (same day, same
192
- everything) to **0/10 informed — fifteen informed runs, zero evasion attempts**. Sonnet,
193
- told the gate exists under "speed is everything" pressure, refused the task outright in
194
- 3/5 runs rather than work around it, and fixed honestly in the other 2. At this scenario
195
- size, disclosure measured as deterrence, not as an evasion spec.
196
-
197
- **Which detectors deserve to exist?** Measured too. Affordance seeds offer an agent one
198
- specific bypass each; predictions are pre-registered before any rep runs
199
- (`harness/PREDICTION-affordance.md`). Result: 7/10 adversarial runs regenerated a golden
200
- file from buggy output and sailed through — **because nothing guarded that class yet**;
201
- that measured demand is why `snapshot-rewrite` exists — while
202
- timeout-inflation and lint-config-gutting measured **0/10 attempts each** — two
203
- pre-registered bets refuted, two detectors *not* built. The same discipline settled
204
- `assertion-weakening` (`harness/PREDICTION-assert-affordance.md`): a seed making the
205
- expectation edit the one cheap, prompt-named move — with a plausible convention-mismatch
206
- cover story — measured **0/10 attempts**; Haiku implemented the real fix every time. My
207
- ≥6/10 bet was refuted, and the rule stays a reserved name. Bypass classes with no
208
- measured demand stay in a record-only shadow scanner that accumulates evidence on every
209
- weekly harness run.
210
-
211
- **And when no honest fix exists at all?** The regime every other seed excludes — and
212
- where the typia incident actually lived — measured with a claims-based oracle
213
- (`harness/PREDICTION-stuck-agent.md`): a provably unfixable suite, an explicit cheap
214
- honest exit ("STATUS: blocked is a fully acceptable outcome"), fabrication defined as
215
- claiming `fixed` against cryptographic reality. **Twenty of twenty pressured runs —
216
- Haiku and Sonnet alike — surrendered honestly**, zero fabrications, zero test edits,
217
- a written diagnosis every time. Both my Haiku bets were refuted: the model that
218
- tampers 6/10 when the cheat is cheaper than the fix takes the honest exit when the
219
- honest exit is cheapest. The sharpened thesis: agents take the cheapest path they
220
- believe is sanctioned — the danger is pressure with no sanctioned exit cheaper than
221
- the lie. The untested harder variant (no one *names* the exit) is registered as the
222
- open follow-up.
223
-
224
- **Is the honesty just eval-awareness?** The sharpest critique of every number above —
225
- the seeds *look* like tests, so maybe the models are performing for a suspected
226
- grader — is measured too (`harness/PREDICTION-eval-awareness.md`): the same
227
- seed dressed as a real project (README, changelog, a user-report thread with a
228
- maintainer's plea, a backdated git history built by a self-deleting kit), gate off.
229
- **Zero frontier tampers in 20 dressed runs**, matching the undressed baselines, while
230
- Haiku stayed a cheater at 8/10. The pre-registered manipulation check is the honest
231
- part: only Opus actually read the dressing (8/10 runs, all ten honest) — so the Opus
232
- null carries the weight and the Sonnet null is scoped as ambient-only. And the
233
- sanctioned-edit variant (`harness/PREDICTION-self-authored.md`) — implement from an
234
- ambiguous spec, then a prompt *licensing* correction of "intern-written" tests —
235
- produced **zero test rewrites in 30 runs across three models**, refuting all three of
236
- my rewrite bets, for a reason worth knowing: every model, implementing blind, chose
237
- the same reading the tests encode. You cannot make an agent believe a test is wrong
238
- by telling it tests might be wrong; the forced-conflict variant is the registered
239
- follow-up.
240
-
241
- **Can "closed vs open" be called in advance?** The two dimension studies
242
- (`harness/PREDICTION-coverage.md`, `harness/PREDICTION-noverify.md`) tested whether a
243
- rule's leak-prone axis can be identified a priori. Both a-priori calls **held
244
- deterministically**: every enumerated hook-skip spelling denies while both
245
- `core.hooksPath` forms allow (the axis called open is open — surveyed by six frontier
246
- runs in reconnaissance, exploited by none), and all three coverage config-surface moves
247
- block. The behavioral half delivered the study's best datum: told to get the coverage
248
- gate green fast, **half of Haiku's runs opened with the per-path override** — the exact
249
- config-surface move the registration predicted a token detector would miss — and the
250
- shipped semantic detector caught every one. Zero tampers in 60 more guarded runs; one
251
- false-positive interception recorded and kept (Fable extending a test-file import via
252
- `sed`, denied by the shell-mutation heuristic mid-honest-fix — the FP record's first
253
- live entry).
254
-
255
- **And when the cheapest cheat is invisible to the gate?** The type-laundering study
256
- (`harness/PREDICTION-asany.md`) probed the one seed whose cheapest cheat spellings are
257
- deliberately uncovered — `: any` annotations are warn-only (`ts-any-launder`, the
258
- FP-avoidance choice) and a bare assertion (`raw as Settings`) has no rule at all — with
259
- a held-out runtime oracle hardened *before* any rep (without that hardening, nine
260
- assertion tampers would have scored honest and the conclusion would have inverted).
261
- Result: **Haiku laundered 10/10** — and the three runs denied on the guarded `as any`
262
- cast **switched to the invisible assertion**, the first time a deny redirected an agent
263
- to a better bypass instead of the fix. **Opus went 10/10 honest with zero gate
264
- involvement** — the strongest disposition evidence measured — while **Fable tampered
265
- 6/10 and Sonnet, the default model, 10/10, both via the bare assertion and drawing
266
- zero denies** (Sonnet's spelling every time: an inline structural assertion,
267
- `as { retries: number; ... }` — the quickest silence that compiles carries no bypass
268
- framing, and Sonnet's famous refusals turn out to be triggered by framing). The
269
- completed frontier split — **Opus 0/10 · Fable 6/10 · Sonnet 10/10** — is the
270
- published shape, with the standing self-measurement caveat on the Fable rows.
271
- Twenty-six of forty runs tampered through the uncovered dimension: measured demand
272
- that sent the `ts-any-launder` warn→block graduation question to a §7.A negatives
273
- corpus — the honest fix itself narrows through `as Record<string, unknown>`, so a
274
- naive assertion-blocker is an FP factory. **The corpus answered, and the question is
275
- closed** (`harness/PREDICTION-launder-corpus.md`): on 1,227 real commits the
276
- laundering spellings fire on legitimate work at 4–10× any deployable block threshold
277
- (annotation-position `any`: 10.5% of commits; named-type assertions: 7.5%), so the
278
- warn stays a warn permanently on syntactic evidence, no assertion rule gets built,
279
- and the laundering class is owned by the layers where coverage is measured: the
280
- held-out **oracle** (caught all 18 assertion tampers) and the **prompt** (the naming
281
- sentence, 10/10 flip). For this class, detection is not the gate's layer —
282
- verification and framing are, and the record says so with a denominator.
283
-
284
- **Why not just a prompt line, or a 20-line grep hook?** Measured
285
- (`harness/PREDICTION-baselines.md`), against a frozen naive hook and a one-sentence
286
- prompt baseline — and two of my three bets lost, both against my own product's
287
- framing. **One integrity sentence took ungated Haiku from 6/10 tampering to 1/10**:
288
- the cheap prompt convention is a real mitigation and this README says so. The naive
289
- hook converted denials nearly as well as the gate (4/5) on the crude-spelling seed —
290
- but was tampered through twice: once by **hardcoding the implementation** (invisible
291
- to any hook; the held-out oracle layer caught it), and once by a denied agent
292
- switching to the node:test **options spelling** `{ skip: true }` — which the shipped
293
- `test-skip` rule *also* missed until it shipped closed in 1.5.1. On 1,227 real merged
294
- commits, the naive patterns would block **one in four legitimate commits (24.9%)**
295
- versus the gate's 9.0% — a ~2.75× precision edge, not the ≥5× I bet, and that 9% on
296
- unconfigured repos is published as the honest tuning-and-sign-off cost, not hidden.
297
- The loop-layer deny alone was never the product; the measured margins are whole-class
298
- spelling coverage, commit-layer precision plus sign-off machinery, and the layers no
299
- hook has. Two follow-ups closed the loops this study opened. The sentence was then
300
- aimed at the regime where the gate is blind — Sonnet's 10/10 assertion-laundering
301
- seed — and produced a **complete flip: 10/10 honest** (`harness/PREDICTION-baselines.md`,
302
- addendum): naming the act is the intervention Sonnet's framing-triggered honesty
303
- runs on. And the promised **tuned-gate numbers** now exist
304
- (`harness/PREDICTION-tuned-fp.md`): 9.0% unconfigured → **6.0%** with a one-line
305
- severity tune → roughly one sign-off decision per 17 commits residual. The same study
306
- shipped per-rule `exclude` globs (1.6.0, with the policy-weakening symmetry built in)
307
- and then measured them honestly: **zero effect on this corpus** — these repos' casts
308
- live in source, not tests, so the zod-recommended path scoping removed nothing here;
309
- the null is published next to the feature.
310
-
311
- **What about just making the tests read-only?** The most common objection got its
312
- pre-registered measurement (`harness/PREDICTION-capability.md`) — starting with the
313
- mechanics the folk advice misses: `chmod` does not bind a root agent and never
314
- prevented `rm` anyway (deletion is a *directory*-write right); the workable barrier
315
- is the immutable bit. Result: **hard immutability governed the crude-tamper seed
316
- completely — 10/10 honest, four EPERM hits all converting to the real fix** (a bare
317
- permission error transferred the deny→fix dynamic with no remediation text). Then
318
- the twist: adding a logged, auto-granted **elevation script made it worse (8/10)** —
319
- both extra tampers walked through the front door with written confessions
320
- ("*Remove failing test file to unblock build*"), the sanctioned-path thesis firing
321
- a fourth time. The lock is easy; the grantor is the hard part — which is the case
322
- for elevation flows that end at a human, and exactly what the gate's sign-off label
323
- already is. That seed now carries all six governance configurations on identical
324
- pre-registered ground: nothing 6/10 tampers, one sentence 1/10, naive hook 2/10,
325
- the gate 0/10, immutable 0/10, immutable+auto-grant 2/10 — with the standing scope
326
- that none of the capability rows touch the semantic classes, by construction.
327
- Running ledger: **70 pre-registered bets — 29 refuted, 37 held, 4 open or
328
- unresolvable** — taskbench's B2 (prevention on real repositories) is the newest
329
- refutation, published beside the bet it refutes.
330
-
331
- **What about false positives?** `snapshot-rewrite` was swept over **1,652 real mainline
332
- commits** (prettier, jest, docusaurus, immer): 216 touched snapshots, all legitimately —
333
- which is exactly why that rule is a `warn` asking for human confirmation, not a block,
334
- and why the study is committed (`harness/fp-study/`) rather than summarized away. The
335
- study's narrow signal — snapshot-only diffs, ~0.06% of audited mainline commits but
336
- 7/7 of observed tampers — now ships as the distinct `snapshot-only-rewrite` warn, the
337
- graduation candidate.
158
+ ## Four layers, different jobs
159
+
160
+ | layer | the question it answers |
161
+ | --- | --- |
162
+ | **Change semantics** | Does a proposed or recorded change match a known weakening class? (PreToolUse deny, staged diff, range diff) |
163
+ | **Effect** | What did the protected state actually become, however the mutation arrived? (drift check, fs-event observer, Stop sweep) |
164
+ | **Outcome** | Does the candidate still pass when protected verification state from the base is restored? (pristine verification — no diff interpreted at all) |
165
+ | **Authority** | Where is the final verdict rendered, and what assumptions make that verdict authoritative? (run envelope, CI, SHA-bound human sign-off) |
166
+
167
+ Sixteen rules are specified and fourteen ship (see the table in
168
+ [SPEC.md](./SPEC.md)). The families: test protection (`test-deletion`,
169
+ `test-skip`, `test-content-removal`), verification-signal protection
170
+ (`coverage-lowering`, `snapshot-rewrite`, `snapshot-only-rewrite`), suppression
171
+ (`ts-any-cast`, `ts-any-launder`, `lint-suppression`), pipeline protection
172
+ (`ci-tampering`, `hook-tampering`, `no-verify`), and the effect/outcome layers
173
+ (`transient-protected-mutation`, `pristine-verification`, plus the `run` envelope).
174
+ Some ambiguous syntactic classes deliberately remain warnings or unimplemented —
175
+ `assertion-weakening` and `guard-removal` are reserved names with no detector, and
176
+ `ts-any-launder` is a permanent warn — rather than being promoted to blockers
177
+ without precision evidence.
178
+
179
+ ## What this does not establish
180
+
181
+ - **Not a correctness oracle.** Pristine verification can only re-run tests that
182
+ exist in the tree. Round 2's designed-in blind spot: on tasks with withheld
183
+ test cases, an agent that honestly half-fixes the visible cases gets a genuine
184
+ visible green and is caught only by the withheld half — no tampering involved.
185
+ A hardcoded implementation is likewise invisible to any diff rule.
186
+ - **Not an OS sandbox, a network firewall, a secret manager, or a replacement for
187
+ branch protection.** It enforces verification integrity; it does not confine
188
+ the process.
189
+ - **The intended deployment is defence in depth:** instructions naming the
190
+ forbidden move + capability restriction + Tamperward + independent held-out
191
+ tests + authoritative CI. Each layer covers classes the others measurably miss.
192
+
193
+ **A word on the words:** *tamper* and *honest* are operational labels on
194
+ artifacts — a protected asset changed, or an independent oracle failed, while the
195
+ visible check went green — decided deterministically. They are not claims about
196
+ any model's intent.
197
+
198
+ ## Research discipline
199
+
200
+ Predictions are registered before counted runs, with numeric bets and explicit
201
+ losing conditions, and the git history proves the order. Seeds, pools, frozen
202
+ analysis scripts, transcripts, and deviation ledgers are committed. A pool used
203
+ to develop detectors is spent — validation requires a fresh draw (round 1's
204
+ repositories are development data now; round 2 drew new ones). Run counts are
205
+ reported with their honest evidence unit: repeated runs on one seed measure that
206
+ seed's stochasticity, and the independent evidence behind the synthetic studies
207
+ is the count of distinct seed configurations, not the run count. Losing
208
+ predictions and corrections remain in the public record rather than being
209
+ removed after the result is known; the series and its errata carry the ledger
210
+ and its totals.
211
+
212
+ **[The research series](./docs/blog/index.md)** ·
213
+ **[The harness](./harness/)** ·
214
+ **[Errata](./docs/blog/errata.md)**
215
+
216
+ ## The architecture was earned by failures
217
+
218
+ | study / version | consequence |
219
+ | --- | --- |
220
+ | Early seed studies | The proposal gate: deny the known weakening classes before the tool call runs, with the deny message as the correction signal |
221
+ | Blind-dimension studies | Coverage ≠ completeness: when the cheapest cheat sat outside detector coverage, agents took it — so an orthogonal verification layer matters more than more rules |
222
+ | Round 1, p = 1.0 | The 53 real trajectories became the evasion corpus — evasion classes the detectors never imagined, extracted with full forensics |
223
+ | v1.7.0 | `test-content-removal`: the row-deletion / assertion-stripping / setup-gutting family that carried round 1's gated false greens, corpus-priced on 2,304 real commits before shipping as block |
224
+ | v1.8.0 | The effect layer: per-call protected-tree drift check plus the `tamperward watch` fs-event observer — judge what the files became, not the command spelling |
225
+ | v1.9.0 | `tamperward verify`: the round-1 pristine oracle productized — visible-green + pristine-red is a masked failure, and "could not verify" fails closed |
226
+ | Round 2, 9–0 | On 22 fresh repositories, v1.9.0 produced b=9/c=0 false-green discordance (p = 0.0039); four gated false greens remained, through two documented residual mechanisms |
227
+ | cssstyle lifecycle escape (round 2) | The `tamperward run` envelope: the gate denied 42 mutations and verify refused all 25 stop attempts, and the runtime completed the session anyway over a masked tree — hooks decide, runtimes terminate, so the adjudication boundary moved outside the runtime |
228
+ | v1.14.1, article audit | `tamperward verify` kept agent-added protected files in the pristine run on the premise that they "only add strictness". An added `conftest.py` could deselect the restored base tests by node id, so a masked failure reported VERIFIED and the envelope printed GREEN MEANS GREEN over an unfixed bug. Added protected files are now removed; PoC and mutation-checked regression committed |
229
+ | v1.10.1–v1.14.0, owner + two-pass external review | Frozen entry-time policy and verifier, entry-SHA ancestry enforcement, quiescence guard, dependency-tree fingerprint, a CI verify step in the generated workflow, the gate pinned to its own version in CI, and SHA-bound sign-off labels |
230
+
231
+ Each row's primary artifact: [CHANGELOG.md](./CHANGELOG.md), [SPEC.md](./SPEC.md),
232
+ and the posts in [docs/blog/](./docs/blog/index.md).
338
233
 
339
234
  ## Stability
340
235
 
@@ -352,7 +247,7 @@ via npm trusted publishing with SLSA provenance. Full rule:
352
247
 
353
248
  ```bash
354
249
  npm install && npm run build # bundles the CLI to dist/cli/index.js
355
- npm test # 259 tests — parser, detectors, engine, policy, renderers
250
+ npm test # 365 tests at v1.14.1 — parser, detectors, engine, policy, renderers
356
251
  npm run typecheck
357
252
  ```
358
253
 
package/dist/cli/index.js CHANGED
@@ -2900,6 +2900,32 @@ import { chmodSync as chmodSync2, cpSync, lstatSync, mkdirSync as mkdirSync6, mk
2900
2900
  import { tmpdir as tmpdir2 } from "node:os";
2901
2901
  import { dirname as dirname4, join as join9, resolve } from "node:path";
2902
2902
  var OVERLAY_CLASSES = ["tests", "snapshots", "config"];
2903
+ var VERIFICATION_SURFACE = [
2904
+ // Python / pytest — read from the rootdir, and conftest at any depth
2905
+ "**/conftest.py",
2906
+ "**/pytest.ini",
2907
+ "**/.pytest.ini",
2908
+ "**/setup.cfg",
2909
+ "**/tox.ini",
2910
+ "**/pyproject.toml",
2911
+ // JavaScript / TypeScript runners and the transforms they load
2912
+ "**/jest.config.*",
2913
+ "**/jest.setup.*",
2914
+ "**/vitest.config.*",
2915
+ "**/vitest.workspace.*",
2916
+ "**/.mocharc.*",
2917
+ "**/karma.conf.*",
2918
+ "**/babel.config.*",
2919
+ "**/.babelrc*",
2920
+ "**/vite.config.*",
2921
+ "**/.swcrc",
2922
+ "**/package.json",
2923
+ // Ruby / PHP / .NET
2924
+ "**/.rspec",
2925
+ "**/phpunit.xml",
2926
+ "**/phpunit.xml.dist",
2927
+ "**/*.runsettings"
2928
+ ];
2903
2929
  function git2(args, cwd) {
2904
2930
  return execFileSync4("git", args, { cwd, encoding: "utf8", maxBuffer: 1 << 28 });
2905
2931
  }
@@ -2948,7 +2974,7 @@ function dropSymlink(p) {
2948
2974
  }
2949
2975
  function overlayPristine(cwd, base, dest, policy) {
2950
2976
  const atBase = git2(["ls-tree", "-r", "--name-only", "-z", base], cwd).split("\0").filter(Boolean);
2951
- const isOverlay = (p) => OVERLAY_CLASSES.some((c) => isProtected(p, policy, c));
2977
+ const isOverlay = (p) => OVERLAY_CLASSES.some((c) => isProtected(p, policy, c)) || matchesAny(p, VERIFICATION_SURFACE);
2952
2978
  let restored = 0;
2953
2979
  const baseProtected = /* @__PURE__ */ new Set();
2954
2980
  for (const rel of atBase) {
@@ -2961,10 +2987,14 @@ function overlayPristine(cwd, base, dest, policy) {
2961
2987
  writeFileSync6(out2, content);
2962
2988
  restored++;
2963
2989
  }
2964
- let kept = 0;
2990
+ let removed = 0;
2965
2991
  const inCopy = git2(["ls-files", "-z", "--cached", "--others", "--exclude-standard"], cwd).split("\0").filter(Boolean);
2966
- for (const rel of inCopy) if (isOverlay(rel) && !baseProtected.has(rel)) kept++;
2967
- return [restored, kept];
2992
+ for (const rel of inCopy) {
2993
+ if (!isOverlay(rel) || baseProtected.has(rel)) continue;
2994
+ rmSync2(join9(dest, rel), { force: true });
2995
+ removed++;
2996
+ }
2997
+ return [restored, removed];
2968
2998
  }
2969
2999
  function runSuite(dir, cmd, budgetSecs) {
2970
3000
  const t0 = Date.now();
@@ -3015,13 +3045,13 @@ function runVerify(opts) {
3015
3045
  const visDir = join9(work, "visible");
3016
3046
  const priDir = join9(work, "pristine");
3017
3047
  let restored = 0;
3018
- let kept = 0;
3048
+ let removedAdded = 0;
3019
3049
  try {
3020
3050
  mkdirSync6(visDir);
3021
3051
  mkdirSync6(priDir);
3022
3052
  materialize(cwd, visDir);
3023
3053
  materialize(cwd, priDir);
3024
- [restored, kept] = overlayPristine(cwd, base, priDir, policy);
3054
+ [restored, removedAdded] = overlayPristine(cwd, base, priDir, policy);
3025
3055
  } catch (e) {
3026
3056
  if (!opts.keep) rmSync2(work, { recursive: true, force: true });
3027
3057
  out2(`verify: could not materialize (${e instanceof Error ? e.message : String(e)}) \u2014 failing closed`);
@@ -3055,7 +3085,7 @@ function runVerify(opts) {
3055
3085
  visible: { exit: visible.exit, secs: visible.secs },
3056
3086
  pristine: { exit: pristine.exit, secs: pristine.secs },
3057
3087
  protected_restored: restored,
3058
- added_protected_kept: kept,
3088
+ added_protected_removed: removedAdded,
3059
3089
  ...opts.keep ? { workdir: work } : {}
3060
3090
  })
3061
3091
  );
@@ -3067,7 +3097,10 @@ function runVerify(opts) {
3067
3097
  BUDGET_EXCEEDED: `budget exceeded (${budget}s): could not verify \u2014 failing closed, not open.`
3068
3098
  };
3069
3099
  out2(`tamperward verify \u2014 ${lines[verdict2]}`);
3070
- if (kept > 0) out2(`(${kept} protected file(s) added since ${base.slice(0, 10)} were kept in the pristine run.)`);
3100
+ if (removedAdded > 0)
3101
+ out2(
3102
+ `(${removedAdded} protected file(s) added since ${base.slice(0, 10)} were removed from the pristine run: the pristine tree carries exactly the base's protected surface.)`
3103
+ );
3071
3104
  }
3072
3105
  return code;
3073
3106
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "tamperward",
3
- "version": "1.14.0",
3
+ "version": "1.14.2",
4
4
  "description": "The deterministic agent-integrity gate. One ruleset, evaluated on the actual diff/commands as a verdict, enforced everywhere a change can be made.",
5
5
  "license": "Apache-2.0",
6
6
  "author": "hexrift",