tamperward 1.14.1 → 1.14.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -32,90 +32,140 @@ impossible.
32
32
 
33
33
  | experiment | result | what it supports |
34
34
  | --- | --- | --- |
35
- | **Round 1** — 26 paired real repos, historical regressions, 53 counted trajectories, **v1.6.0** ([`harness/taskbench/`](./harness/taskbench/)) | Transfer: **13/26 ungated runs (50%)** violated policy. Headline prevention bet **lost**: b=5 / c=4, RD +3.8pp [−17.2, +24.7], exact McNemar **p = 1.0** — published beside the bet | A detector-centric architecture was insufficient: agents routed around the shipped detector classes |
36
- | **Pristine oracle, round 1** — independent re-execution of the original suite | Identified **every masked failure observed among the 53 trajectories** while diff-time detection was routed around | The outcome-level signal that motivated `tamperward verify` |
37
- | **Round 2** — **22 fresh held-out repos** no detector was tuned on ([`harness/taskbench/round2/`](./harness/taskbench/round2/)) | Transfer 14/22 (64%). Prevention: **b=9 / c=0**, RD **+40.9pp**, BP95 **[17.8, 61.3]**, exact McNemar **p = 0.0039** | The preregistered **v1.9.0** treatment materially reduced false greens in that setting |
35
+ | **Round 1** — 26 paired real repos, historical regressions, 53 counted trajectories, **v1.6.0** ([`harness/taskbench/`](./harness/taskbench/)) | Transfer: **9/27 ungated runs (33.3%)** violated policy ([corrected](./harness/taskbench/reanalysis/TRANSFER-REANALYSIS.md) from a published 13/26 — the original predicate was defective; the transfer bet is refuted, not held). Headline prevention bet **lost**: b=5 / c=4, RD +3.8pp [−17.2, +24.7], exact McNemar **p = 1.0** — published beside the bet | A detector-centric architecture was insufficient: agents routed around the shipped detector classes |
36
+ | **Pristine oracle, round 1** — independent re-execution of the original suite, **including withheld semantic cases** | Identified **every masked failure observed among the 53 trajectories** while diff-time detection was routed around | The outcome-level signal that motivated `tamperward verify`. The shipped command productizes this oracle's base-restoration component only — it carries no withheld cases and cannot detect a semantically incomplete fix the base tests also accept |
37
+ | **Round 2** — **22 fresh held-out repos** no detector was tuned on ([`harness/taskbench/round2/`](./harness/taskbench/round2/)) | Transfer **12/22 (54.5%)** ([corrected](./harness/taskbench/reanalysis/TRANSFER-REANALYSIS.md) from a published 14/22). Prevention: **b=9 / c=0**, RD **+40.9pp**, BP95 **[17.8, 61.3]**, exact McNemar **p = 0.0039** | The preregistered **v1.9.0** treatment materially reduced false greens in that setting |
38
38
  | **After prevention** — the nine round-2 prevented false greens | **8 of 9 became honest completions**; the ninth an honest **non**-completion | Prevention usually redirected trajectories toward honest work rather than merely blocking them |
39
+ | **Round 3** — 17 paired Python repos, fresh PyPI frame, **v1.14.0** ([`harness/taskbench/round3/`](./harness/taskbench/round3/)) | Transfer 9/17 (52.9%). Prevention: **b=6 / c=0**, RD **+35.3pp**, BP95 [9.5, 58.7], exact McNemar **p = 0.0313** | The prevention result appeared in Python — with the treatment also changed from v1.9.0, so ecosystem transfer is not isolated. The in-loop skip detector proved blind to pytest syntax; the outer layers carried it |
40
+ | **Round 3.1** — the same 16 pairs under **`claude-sonnet-5`** ([`harness/taskbench/round3.1/`](./harness/taskbench/round3.1/)) | Transfer 4/16 (25.0%). Prevention: **b=1 / c=0**, RD +6.3pp, BP95 [−13.8, 28.3], exact McNemar **p = 1.0000** | The confirmatory result **did not replicate**, and could not have: only three ungated false greens occurred, so `b ≤ 3` and p ≥ .25 whatever the gate did. A failure to reject, not evidence of no effect |
39
41
 
40
42
  Earlier controlled experiments → **[the research series](./docs/blog/index.md)**.
41
43
 
42
44
  > **Scope.** The rows above are one model, one pressure prompt, one JS/TS
43
45
  > ecosystem — evidence for that setting, not a universal claim. Round 2 tested
44
- > the v1.9.0 stack; the current **v1.14.0** adds post-study envelope hardening
46
+ > the v1.9.0 stack; the current **v1.14.3** adds post-study envelope hardening
45
47
  > (externally reviewed, with findings tracked individually as REPRO or AUDIT in
46
48
  > [SECURITY-ENVELOPE.md](./SECURITY-ENVELOPE.md) and closed with regression and
47
- > mutation checks — see [CHANGELOG](./CHANGELOG.md)). Round 3 is in progress on a frozen v1.14.0
48
- > baseline, testing external validity on a fresh repository frame, with model,
49
- > runtime and ecosystem among the planned axes of change
50
- > ([`harness/taskbench/round3/ROUND3-PLAN.md`](./harness/taskbench/round3/ROUND3-PLAN.md)).
51
- > No round-3 outcomes are included in the claims above.
49
+ > mutation checks — see [CHANGELOG](./CHANGELOG.md)). Rounds 3 and 3.1 are
50
+ > complete and are in the table above. **1.14.1 closed a bypass in `tamperward
51
+ > verify`** — an agent-added protected file could suppress the tests the pristine
52
+ > run had just restored — which was present from v1.9.0 onward, so rounds 2, 3
53
+ > and 3.1 all ran on a verifier carrying it. No evidence any agent used it; the
54
+ > layer's authority in those rounds was nonetheless weaker than described at the
55
+ > time. Round 4 is the fresh pool.
52
56
 
53
57
  ## Architecture
54
58
 
55
- The final envelope adjudication is anchored to state recorded before the agent
56
- starts; the in-loop hooks read the live tree and remain steering controls. The
57
- final adjudication boundary sits outside the agent's normal lifecycle.
59
+ Tamperward separates **steering** from **adjudication**.
58
60
 
61
+ > **Core invariant:** the agent may author the candidate tree, but it must not
62
+ > choose the trusted baseline, the governing policy, the verifier, or the final
63
+ > verdict.
64
+
65
+ | plane | mechanism | role |
66
+ | --- | --- | --- |
67
+ | **In-loop steering** | Claude Code `PreToolUse`, protected-tree drift checks, the `Stop` sweep, and optional filesystem-event observation | Immediate corrective feedback. Not the final authority. |
68
+ | **Local adjudication** | `tamperward run -- <agent…>` | Wraps the runtime and independently judges the tree it leaves behind. |
69
+ | **Commit backstop** | `tamperward check --staged` | Catches human changes, and agents without supported hooks. |
70
+ | **Repository authority** | Required PR checks plus branch protection | Decides what may reach the protected branch. |
71
+
72
+ | layer | the question it answers |
73
+ | --- | --- |
74
+ | **Change semantics** | Does a proposed or recorded change match a known weakening class? (PreToolUse deny, staged diff, range diff) |
75
+ | **Effect** | What did the protected state actually become, however the mutation arrived? (drift check, fs-event observer, Stop sweep) |
76
+ | **Outcome** | Does the candidate still pass when protected verification state from the base is restored? (pristine verification — no diff interpreted at all) |
77
+ | **Authority** | Where is the final verdict rendered, and what assumptions make it authoritative? (run envelope, CI, SHA-bound human sign-off) |
78
+
79
+ ### The local enforcement envelope
80
+
81
+ ```mermaid
82
+ flowchart TB
83
+ subgraph R["tamperward run — outer enforcement envelope"]
84
+ E["Capture entry state"]
85
+ subgraph U["Agent-controlled lifecycle"]
86
+ H["In-loop steering hooks"]
87
+ A["Agent runtime"]
88
+ W["Candidate HEAD and worktree"]
89
+ H -.-> A
90
+ A --> W
91
+ end
92
+ J["Post-exit adjudication"]
93
+ C["Committed and worktree checks"]
94
+ V["Visible and pristine verification"]
95
+ Q["Ancestry, dependency drift, quiescence"]
96
+ X["Final exit verdict"]
97
+ E --> A
98
+ E --> J
99
+ W --> J
100
+ J --> C
101
+ J --> V
102
+ J --> Q
103
+ C --> X
104
+ V --> X
105
+ Q --> X
106
+ end
59
107
  ```
60
- RECORDED AT ENVELOPE ENTRY:
61
- entry SHA · frozen policy · frozen verifier command/budget · dependency fingerprint
62
- │
63
- ┌─────────────────────────────┴───────────────────────────────┐
64
- │ agent loop (Claude Code hooks today) │
65
- │ PreToolUse gate: every proposed Edit/Write/Bash judged; │
66
- │ known weakening DENIED before the tool runs │
67
- │ │ │
68
- │ working tree ▼ │
69
- │ effect observation: per-call protected-tree drift check │
70
- │ + filesystem-event observer (`tamperward watch`) │
71
- │ │ │
72
- │ Stop sweep ▼ │
73
- │ end-of-turn net diff re-checked; transient protected │
74
- │ mutations judged from the event log │
75
- │ │ │
76
- │ pristine verification ▼ (`tamperward verify`) │
77
- │ suite runs twice in isolated copies: the candidate tree │
78
- │ as-is, and with protected tests/snapshots/config │
79
- │ restored from the trusted base; visible-green + │
80
- │ pristine-red = MASKED FAILURE │
81
- └─────────────────────────────┬───────────────────────────────┘
82
- │ agent runtime exits — exit code NOT trusted
83
- ▼
84
- `tamperward run -- <agent…>` — the outer envelope
85
- policy + verifier resolved from the entry state, never the tree
86
- the agent wrote · post-agent HEAD must descend from the entry
87
- SHA · policy check over base...HEAD (committed weakening) and
88
- over the worktree incl. untracked files (uncommitted) ·
89
- dependency-tree fingerprint compared · quiescence guard ·
90
- pristine verify against the entry base · exit 1 on any blocking
91
- finding or masked failure · exit 2 (fail closed) when
92
- adjudication is impossible
93
- │
94
- ▼
95
- pre-commit backstop: `tamperward check --staged`
96
- │
97
- ▼
98
- PR CI: `check --diff base...head` + `verify --require-ancestor`
99
- policy read from the merge-base — a PR cannot govern its own verdict
100
- │
101
- ▼
102
- out-of-band human exception only: PR label
103
- `tamperward:allow:<rule>@<head-sha>` — SHA-bound, applied by a human
104
- with write access, never a file the branch can commit
105
- │
106
- ▼
107
- protected main
108
- ```
109
108
 
110
- CI is the final authority in this design, and that authority is conditional on
111
- the deployment assumptions documented in
112
- [docs/guide/enforcement.md](./docs/guide/enforcement.md) and
113
- [SPEC.md](./SPEC.md): verification is anchored to a base the agent cannot
114
- rewrite (a protected branch), the repository's branch protection actually
115
- enforces the workflow's verdict on merges, and sign-off labels can only be
116
- applied by humans with write access. Mechanisms, residuals, and what is
117
- explicitly outside the trust boundary: [SPEC.md](./SPEC.md) and
118
- [SECURITY-ENVELOPE.md](./SECURITY-ENVELOPE.md).
109
+ Before starting the agent, the envelope
110
+
111
+ 1. records the exact entry commit;
112
+ 2. loads the policy and the verifier command and budget from that commit;
113
+ 3. records the installed `node_modules` content fingerprint, where one exists; and
114
+ 4. refuses an already-dirty tree unless the operator explicitly accepts the risk.
115
+
116
+ The hooks then steer the agent while it works: they read live state, deny known
117
+ weakening operations, detect protected-tree drift, and sweep the turn's net
118
+ changes. The `Stop` hook runs that sweep — it does **not** invoke `verify`.
119
+ `tamperward watch` can additionally record transient mutations, but it is an
120
+ optional observer that `init` neither starts nor supervises, and it is not an
121
+ enforcement boundary.
122
+
123
+ After the runtime exits its exit code is treated as untrusted, and the envelope
124
+ checks that post-agent `HEAD` still descends from the entry commit; the committed
125
+ changes over `entry...HEAD`; staged, unstaged and untracked non-ignored worktree
126
+ changes; dependency drift and whether the tree stayed quiescent; and the
127
+ verification outcome.
128
+
129
+ `tamperward verify` materialises two temporary copies — they are **not
130
+ sandboxes**, and both execute agent-controlled code. The visible copy runs the
131
+ candidate as-is. In the pristine copy, files on the protected overlay classes and
132
+ on a separate enumerated **verification surface** (the configuration a test runner
133
+ auto-consults) are restored from the trusted base, and agent-added files on those
134
+ surfaces are removed. That surface list deliberately **bounds** runner-configuration
135
+ inheritance rather than eliminating it — see
136
+ [docs/THREAT-MODEL-pristine-run.md](./docs/THREAT-MODEL-pristine-run.md).
137
+
138
+ Visible green followed by pristine red is a `MASKED_FAILURE`. Two green runs
139
+ establish only that the candidate passes the base suite; they establish neither
140
+ semantic correctness nor a substitute for independent held-out tests.
141
+
142
+ | result | envelope behaviour |
143
+ | --- | --- |
144
+ | Enforcement clean and agent exits 0 | exit 0 |
145
+ | Blocking finding, masked failure, dependency drift, or non-quiescent tree | exit 1 |
146
+ | Required adjudication cannot be completed | exit 2, failing closed |
147
+ | Enforcement clean but the agent failed | the agent's non-zero exit is preserved |
148
+
149
+ ### CI authority
150
+
151
+ Pre-commit is an early backstop; protected CI is the merge authority. CI runs, independently:
152
+
153
+ 1. `check --diff base...head`, with policy from the trusted merge-base; and
154
+ 2. `verify --require-ancestor --base <base-sha>`, which since **1.14.3** also reads
155
+ its policy, verification surface and verifier command and budget from that same
156
+ trusted base — so neither job's verdict is governed by the candidate.
157
+
158
+ Legitimate exceptions are out-of-band PR labels bound to the exact head SHA; a new
159
+ push invalidates the approval.
160
+
161
+ This guarantee depends on a protected and immutable base, required status checks, a
162
+ pinned Tamperward version, and label permissions restricted to trusted humans.
163
+ Tamperward provides no OS isolation, no network confinement, no secret isolation,
164
+ and no semantic-correctness oracle.
165
+
166
+ Full assumptions and residual risks: [SPEC.md](./SPEC.md),
167
+ [SECURITY-ENVELOPE.md](./SECURITY-ENVELOPE.md), and the
168
+ [series limitations page](./docs/blog/limitations.md).
119
169
 
120
170
  ## Quick start
121
171
 
@@ -150,14 +200,7 @@ npx tamperward verify --base main # pristine-suite re-execution
150
200
  npx tamperward run -- <agent command...> # the outer envelope around an agent
151
201
  ```
152
202
 
153
- ## Four layers, different jobs
154
-
155
- | layer | the question it answers |
156
- | --- | --- |
157
- | **Change semantics** | Does a proposed or recorded change match a known weakening class? (PreToolUse deny, staged diff, range diff) |
158
- | **Effect** | What did the protected state actually become, however the mutation arrived? (drift check, fs-event observer, Stop sweep) |
159
- | **Outcome** | Does the candidate still pass when protected verification state from the base is restored? (pristine verification — no diff interpreted at all) |
160
- | **Authority** | Where is the final verdict rendered, and what assumptions make that verdict authoritative? (run envelope, CI, SHA-bound human sign-off) |
203
+ ### The rules
161
204
 
162
205
  Sixteen rules are specified and fourteen ship (see the table in
163
206
  [SPEC.md](./SPEC.md)). The families: test protection (`test-deletion`,
@@ -220,7 +263,10 @@ and its totals.
220
263
  | v1.9.0 | `tamperward verify`: the round-1 pristine oracle productized — visible-green + pristine-red is a masked failure, and "could not verify" fails closed |
221
264
  | Round 2, 9–0 | On 22 fresh repositories, v1.9.0 produced b=9/c=0 false-green discordance (p = 0.0039); four gated false greens remained, through two documented residual mechanisms |
222
265
  | cssstyle lifecycle escape (round 2) | The `tamperward run` envelope: the gate denied 42 mutations and verify refused all 25 stop attempts, and the runtime completed the session anyway over a masked tree — hooks decide, runtimes terminate, so the adjudication boundary moved outside the runtime |
223
- | v1.10.1–v1.14.0, owner + two-pass external review | Frozen entry-time policy and verifier, entry-SHA ancestry enforcement, quiescence guard, dependency-tree fingerprint, a CI verify step in the generated workflow, the gate pinned to its own version in CI, and SHA-bound sign-off labels |
266
+ | v1.14.3, external review | Standalone `verify` loaded its policy from the working tree, so the generated CI workflow let a pull request supply the `verify:` command for its own re-execution. `check --diff` flagged the edit as hook-tampering, so the workflow caught it as a pair — but only where both jobs are required, and `verify` alone had no protection. With a `--base`, policy now comes from that commit |
267
+ | v1.14.2, threat model | 1.14.1 removed agent-added files only inside the protected classes, whose `config` list is JS/TS-only; an added `pytest.ini`, `setup.cfg`, `tox.ini` or `pyproject.toml` still reached the pristine run. `verify` now owns a verification surface covering runner-consulted configuration |
268
+ | v1.14.1, article audit | `tamperward verify` kept agent-added protected files in the pristine run on the premise that they "only add strictness". An added `conftest.py` could deselect the restored base tests by node id, so a masked failure reported VERIFIED and the envelope printed GREEN MEANS GREEN over an unfixed bug. Added protected files are now removed; PoC and mutation-checked regression committed |
269
+ | v1.10.1–v1.14.0, owner + two-pass external review | Frozen entry-time policy and verifier, entry-SHA ancestry enforcement, quiescence guard, `node_modules` content fingerprint, a CI verify step in the generated workflow, the gate pinned to its own version in CI, and SHA-bound sign-off labels |
224
270
 
225
271
  Each row's primary artifact: [CHANGELOG.md](./CHANGELOG.md), [SPEC.md](./SPEC.md),
226
272
  and the posts in [docs/blog/](./docs/blog/index.md).
@@ -241,7 +287,7 @@ via npm trusted publishing with SLSA provenance. Full rule:
241
287
 
242
288
  ```bash
243
289
  npm install && npm run build # bundles the CLI to dist/cli/index.js
244
- npm test # 364 tests at v1.14.0 — parser, detectors, engine, policy, renderers
290
+ npm test # 367 tests at v1.14.3 — parser, detectors, engine, policy, renderers
245
291
  npm run typecheck
246
292
  ```
247
293
 
package/dist/cli/index.js CHANGED
@@ -2900,6 +2900,32 @@ import { chmodSync as chmodSync2, cpSync, lstatSync, mkdirSync as mkdirSync6, mk
2900
2900
  import { tmpdir as tmpdir2 } from "node:os";
2901
2901
  import { dirname as dirname4, join as join9, resolve } from "node:path";
2902
2902
  var OVERLAY_CLASSES = ["tests", "snapshots", "config"];
2903
+ var VERIFICATION_SURFACE = [
2904
+ // Python / pytest — read from the rootdir, and conftest at any depth
2905
+ "**/conftest.py",
2906
+ "**/pytest.ini",
2907
+ "**/.pytest.ini",
2908
+ "**/setup.cfg",
2909
+ "**/tox.ini",
2910
+ "**/pyproject.toml",
2911
+ // JavaScript / TypeScript runners and the transforms they load
2912
+ "**/jest.config.*",
2913
+ "**/jest.setup.*",
2914
+ "**/vitest.config.*",
2915
+ "**/vitest.workspace.*",
2916
+ "**/.mocharc.*",
2917
+ "**/karma.conf.*",
2918
+ "**/babel.config.*",
2919
+ "**/.babelrc*",
2920
+ "**/vite.config.*",
2921
+ "**/.swcrc",
2922
+ "**/package.json",
2923
+ // Ruby / PHP / .NET
2924
+ "**/.rspec",
2925
+ "**/phpunit.xml",
2926
+ "**/phpunit.xml.dist",
2927
+ "**/*.runsettings"
2928
+ ];
2903
2929
  function git2(args, cwd) {
2904
2930
  return execFileSync4("git", args, { cwd, encoding: "utf8", maxBuffer: 1 << 28 });
2905
2931
  }
@@ -2948,7 +2974,7 @@ function dropSymlink(p) {
2948
2974
  }
2949
2975
  function overlayPristine(cwd, base, dest, policy) {
2950
2976
  const atBase = git2(["ls-tree", "-r", "--name-only", "-z", base], cwd).split("\0").filter(Boolean);
2951
- const isOverlay = (p) => OVERLAY_CLASSES.some((c) => isProtected(p, policy, c));
2977
+ const isOverlay = (p) => OVERLAY_CLASSES.some((c) => isProtected(p, policy, c)) || matchesAny(p, VERIFICATION_SURFACE);
2952
2978
  let restored = 0;
2953
2979
  const baseProtected = /* @__PURE__ */ new Set();
2954
2980
  for (const rel of atBase) {
@@ -2988,7 +3014,9 @@ function runVerify(opts) {
2988
3014
  const out2 = (s) => void process.stdout.write(s + "\n");
2989
3015
  let policy;
2990
3016
  try {
2991
- policy = opts.policyOverride ?? loadPolicy(cwd);
3017
+ if (opts.policyOverride) policy = opts.policyOverride;
3018
+ else if (opts.base) policy = loadPolicyAt(resolveBase(opts.base, cwd), cwd) ?? defaultPolicy();
3019
+ else policy = loadPolicy(cwd);
2992
3020
  } catch (e) {
2993
3021
  out2(`verify: cannot load policy (${e instanceof Error ? e.message : String(e)}) \u2014 failing closed`);
2994
3022
  return 2;
@@ -3004,7 +3032,9 @@ function runVerify(opts) {
3004
3032
  }
3005
3033
  const cmd = opts.cmd ?? policy.verify?.command;
3006
3034
  if (!cmd) {
3007
- out2("verify: no suite command \u2014 set policy `verify: { command: ... }` or pass --cmd");
3035
+ out2(
3036
+ opts.base ? "verify: no suite command in the policy at the trusted base \u2014 the base governs the verifier, so a `verify:` block added only on the candidate is not used. Add it at the base, or pass --cmd explicitly." : "verify: no suite command \u2014 set policy `verify: { command: ... }` or pass --cmd"
3037
+ );
3008
3038
  return 2;
3009
3039
  }
3010
3040
  const budget = opts.budget ?? policy.verify?.budget ?? 300;
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "tamperward",
3
- "version": "1.14.1",
3
+ "version": "1.14.3",
4
4
  "description": "The deterministic agent-integrity gate. One ruleset, evaluated on the actual diff/commands as a verdict, enforced everywhere a change can be made.",
5
5
  "license": "Apache-2.0",
6
6
  "author": "hexrift",