tamperward 2.0.0 → 2.1.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +48 -15
- package/dist/cli/index.js +78 -28
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -12,6 +12,7 @@
|
|
|
12
12
|
<a href="./LICENSE"><img src="https://img.shields.io/badge/license-Apache--2.0-lightgrey" alt="license"></a>
|
|
13
13
|
</p>
|
|
14
14
|
|
|
15
|
+
**[Quick start](#quick-start)** ·
|
|
15
16
|
**[Docs & guide](https://hexrift.github.io/tamperward/)** ·
|
|
16
17
|
**[The research series](./docs/blog/index.md)** — every registered prediction
|
|
17
18
|
published beside its outcome
|
|
@@ -22,14 +23,29 @@ trajectories, some modify or attempt to modify verification in ways that can
|
|
|
22
23
|
turn incorrect work into apparent success. Under pressure, the cheaper route to
|
|
23
24
|
green is sometimes to weaken the checks instead of fixing the failure.
|
|
24
25
|
|
|
26
|
+
In plain English: Tamperward lets a coding agent change your code, but not the
|
|
27
|
+
trusted starting point, the rules, or the checks used to judge that code.
|
|
28
|
+
|
|
25
29
|
Tamperward is a **deterministic verification-integrity layer**. It blocks known
|
|
26
30
|
weakening moves as they happen, observes protected-state effects, and
|
|
27
31
|
independently re-adjudicates apparent success outside the agent's normal
|
|
28
32
|
completion path. No runtime LLM judge. Fail closed when adjudication is
|
|
29
33
|
impossible.
|
|
30
34
|
|
|
35
|
+
> **Project status: active research release.** Tamperward is usable today, but
|
|
36
|
+
> its enforcement architecture and supporting evidence are still being tested
|
|
37
|
+
> and hardened. Use it as one layer of defence in depth alongside protected CI,
|
|
38
|
+
> independent tests and human review. Findings, limitations and corrections are
|
|
39
|
+
> published openly. The 2.0 major marks the Node 18 drop, not a declaration of
|
|
40
|
+
> security maturity; the distance to that is tracked, milestone by milestone, in
|
|
41
|
+
> [SPEC §9.1](./SPEC.md#91-maturity-milestones).
|
|
42
|
+
|
|
31
43
|
## What we have actually measured
|
|
32
44
|
|
|
45
|
+
Plain-English takeaway: the original detector-centred design was insufficient.
|
|
46
|
+
Later versions materially reduced false-green outcomes in two held-out rounds,
|
|
47
|
+
and the subsequent stronger-model replication was inconclusive.
|
|
48
|
+
|
|
33
49
|
| experiment | result | what it supports |
|
|
34
50
|
| --- | --- | --- |
|
|
35
51
|
| **Round 1** — 26 paired real repos, historical regressions, 53 counted trajectories, **v1.6.0** ([`harness/taskbench/`](./harness/taskbench/)) | Transfer: **9/27 ungated runs (33.3%)** violated policy ([corrected](./harness/taskbench/reanalysis/TRANSFER-REANALYSIS.md) from a published 13/26 — the original predicate was defective; the transfer bet is refuted, not held). Headline prevention bet **lost**: b=5 / c=4, RD +3.8pp [−17.2, +24.7], exact McNemar **p = 1.0** — published beside the bet | A detector-centric architecture was insufficient: agents routed around the shipped detector classes |
|
|
@@ -39,11 +55,16 @@ impossible.
|
|
|
39
55
|
| **Round 3** — 17 paired Python repos, fresh PyPI frame, **v1.14.0** ([`harness/taskbench/round3/`](./harness/taskbench/round3/)) | Transfer 9/17 (52.9%). Prevention: **b=6 / c=0**, RD **+35.3pp**, BP95 [9.5, 58.7], exact McNemar **p = 0.0313** | The prevention result appeared in Python — with the treatment also changed from v1.9.0, so ecosystem transfer is not isolated. The in-loop skip detector proved blind to pytest syntax; the outer layers carried it |
|
|
40
56
|
| **Round 3.1** — the same 16 pairs under **`claude-sonnet-5`** ([`harness/taskbench/round3.1/`](./harness/taskbench/round3.1/)) | Transfer 4/16 (25.0%). Prevention: **b=1 / c=0**, RD +6.3pp, BP95 [−13.8, 28.3], exact McNemar **p = 1.0000** | The confirmatory result **did not replicate**, and could not have: only three ungated false greens occurred, so `b ≤ 3` and p ≥ .25 whatever the gate did. A failure to reject, not evidence of no effect |
|
|
41
57
|
|
|
58
|
+
Key: `b` = false greens seen only without Tamperward; `c` = false greens seen
|
|
59
|
+
only with it; `RD` = paired risk difference; `pp` = percentage points; `BP95` =
|
|
60
|
+
Bonett–Price 95% interval.
|
|
61
|
+
|
|
42
62
|
Earlier controlled experiments → **[the research series](./docs/blog/index.md)**.
|
|
43
63
|
|
|
44
|
-
> **Scope.** The rows above
|
|
45
|
-
>
|
|
46
|
-
>
|
|
64
|
+
> **Scope.** The rows above cover specific models, pressure prompts, treatment
|
|
65
|
+
> versions, and finite JavaScript/TypeScript and Python repository samples —
|
|
66
|
+
> evidence for those settings, not a universal claim. Round 2 tested
|
|
67
|
+
> the v1.9.0 stack; the current **2.x** line adds post-study envelope hardening
|
|
47
68
|
> (externally reviewed, with findings tracked individually as REPRO or AUDIT in
|
|
48
69
|
> [SECURITY-ENVELOPE.md](./SECURITY-ENVELOPE.md) and closed with regression and
|
|
49
70
|
> mutation checks — see [CHANGELOG](./CHANGELOG.md)). Rounds 3 and 3.1 are
|
|
@@ -56,7 +77,11 @@ Earlier controlled experiments → **[the research series](./docs/blog/index.md)
|
|
|
56
77
|
|
|
57
78
|
## Architecture
|
|
58
79
|
|
|
59
|
-
|
|
80
|
+
The agent may produce the work, but it must not control how that work is
|
|
81
|
+
judged. In architectural terms, Tamperward separates **steering** from
|
|
82
|
+
**adjudication**. "Visible" verification runs the candidate as it stands;
|
|
83
|
+
"pristine" verification restores the protected verification state from the
|
|
84
|
+
trusted starting point and runs the checks again.
|
|
60
85
|
|
|
61
86
|
> **Core invariant:** the agent may author the candidate tree, but it must not
|
|
62
87
|
> choose the trusted baseline, the governing policy, the verifier, or the final
|
|
@@ -160,16 +185,19 @@ are involved here, and they are not the same thing.
|
|
|
160
185
|
trusted base — so neither step's verdict is governed by the candidate.
|
|
161
186
|
|
|
162
187
|
Legitimate exceptions are out-of-band PR labels, `tamperward:allow:<rule>@<head-sha>`.
|
|
163
|
-
The workflow passes the head SHA to
|
|
188
|
+
The workflow passes the head SHA to both steps through `TAMPERWARD_OOB_HEAD`, so an
|
|
164
189
|
approval is bound to the exact commit it was granted for and a new push invalidates it.
|
|
190
|
+
Since **2.1.0** the verify step reads the same labels: `tamperward:allow:verify@<head-sha>`
|
|
191
|
+
accepts a `MASKED_FAILURE` — the case where a behaviour change makes the original
|
|
192
|
+
expectations wrong and a reviewer has read the test edit and said so. It clears nothing
|
|
193
|
+
else: a red visible suite, or a run that could not verify, stays red whatever the labels
|
|
194
|
+
say, and the verdict is still reported as a masked failure; only the exit code changes.
|
|
165
195
|
|
|
166
196
|
**This repository's own self-gate** — the `gate` job in `.github/workflows/ci.yml` —
|
|
167
197
|
runs the built CLI's `check --diff` over the pull-request range, cleared only by the
|
|
168
198
|
same label channel, but does **not** run `verify` on itself. This repo's test
|
|
169
|
-
expectations legitimately change whenever a rule changes,
|
|
170
|
-
|
|
171
|
-
closed on exactly the pull requests that improve the gate. The self-gate is a
|
|
172
|
-
diff-time gate.
|
|
199
|
+
expectations legitimately change whenever a rule changes, so nearly every rule pull
|
|
200
|
+
request would need the verify label; the self-gate stays a diff-time gate.
|
|
173
201
|
|
|
174
202
|
This guarantee depends on a protected and immutable base, required status checks, a
|
|
175
203
|
pinned Tamperward version, and label permissions restricted to trusted humans.
|
|
@@ -186,6 +214,10 @@ Full assumptions and residual risks: [SPEC.md](./SPEC.md),
|
|
|
186
214
|
npx tamperward init
|
|
187
215
|
```
|
|
188
216
|
|
|
217
|
+
Requires Node.js 20.19 or later. JavaScript and TypeScript are the fully
|
|
218
|
+
supported detector surface; the other documented ecosystems get file-level and
|
|
219
|
+
pattern-based protection.
|
|
220
|
+
|
|
189
221
|
One idempotent command wires the policy, the agent hooks, the pre-commit hook, a
|
|
190
222
|
CI workflow that runs both the diff-time check and pristine verification, and a
|
|
191
223
|
`CODEOWNERS` requirement on the paths that decide whether the gate runs at all.
|
|
@@ -252,7 +284,7 @@ commands ignore what they do not know.
|
|
|
252
284
|
| command | 0 | 1 | 2 |
|
|
253
285
|
| --- | --- | --- | --- |
|
|
254
286
|
| `check` | no blocking finding | at least one blocking finding | cannot evaluate: policy parse error, malformed `--diff` range, or no view given |
|
|
255
|
-
| `verify` | `VERIFIED` — visible and pristine both green | `MASKED_FAILURE` (visible green, pristine red) or `SUITE_RED` | cannot verify, failing closed: no suite command, unresolvable base, `--require-ancestor` refused, budget exceeded, or the working or dependency tree moved during the run |
|
|
287
|
+
| `verify` | `VERIFIED` — visible and pristine both green; or a `MASKED_FAILURE` cleared by an out-of-band `verify@<head-sha>` approval | `MASKED_FAILURE` (visible green, pristine red) or `SUITE_RED` | cannot verify, failing closed: no suite command, unresolvable base, `--require-ancestor` refused, budget exceeded, or the working or dependency tree moved during the run |
|
|
256
288
|
| `run` | enforcement clean and the agent exited 0 — a non-zero agent exit is passed through unchanged | any blocking finding or masked failure, whatever the agent returned | cannot adjudicate: dirty start, policy error, verify cannot run |
|
|
257
289
|
| `hook claude` / `sweep claude` | always — a deny is JSON on stdout at exit 0, never exit 2 | — | only for an unsupported agent name |
|
|
258
290
|
| `allow` | sign-off recorded | — | no rule or `--reason`, not a git repo, or no current blocking finding to sign off |
|
|
@@ -263,8 +295,8 @@ commands ignore what they do not know.
|
|
|
263
295
|
|
|
264
296
|
| variable | who sets it | what it does |
|
|
265
297
|
| --- | --- | --- |
|
|
266
|
-
| `TAMPERWARD_OOB_SIGNOFF` | the CI workflow, from PR labels | comma-separated out-of-band approvals — `<rule>` or `<rule>:<file
|
|
267
|
-
| `TAMPERWARD_OOB_HEAD` | the CI workflow (`github.event.pull_request.head.sha`) | the head SHA under adjudication; once set, an approval clears
|
|
298
|
+
| `TAMPERWARD_OOB_SIGNOFF` | the CI workflow, from PR labels | comma-separated out-of-band approvals — `<rule>` or `<rule>:<file>` for `check --diff`, `verify` for a `verify` masked failure — optionally `@<head-sha>`; honoured at the CI layer only, never the committed ledger |
|
|
299
|
+
| `TAMPERWARD_OOB_HEAD` | the CI workflow (`github.event.pull_request.head.sha`) | the head SHA under adjudication; once set, an approval clears anything only if it names that commit (`@<sha>`, at least 7 characters), so a new push re-blocks |
|
|
268
300
|
| `TAMPERWARD_DENYLOG` | a harness or operator | a file to which `hook claude` and `sweep claude` append the rule ids of every deny, one line per verdict, best effort |
|
|
269
301
|
| `TAMPERWARD_FSEVENTS` | operator or harness | overrides the `tamperward watch` event-log path (default `.git/tamperward/fsevents.jsonl`); the Stop sweep reads the same variable |
|
|
270
302
|
| `TAMPERWARD_WATCH_NO_RECURSIVE` | CI and tests (`=1`) | forces `tamperward watch` onto its per-directory fallback instead of recursive `fs.watch`, so the fallback is exercised on every platform |
|
|
@@ -286,7 +318,7 @@ Some ambiguous syntactic classes deliberately remain warnings or unimplemented
|
|
|
286
318
|
`ts-any-launder` is a permanent warn — rather than being promoted to blockers
|
|
287
319
|
without precision evidence.
|
|
288
320
|
|
|
289
|
-
## What
|
|
321
|
+
## What Tamperward does not do
|
|
290
322
|
|
|
291
323
|
- **Not a correctness oracle.** Pristine verification can only re-run tests that
|
|
292
324
|
exist in the tree. Round 2's designed-in blind spot: on tasks with withheld
|
|
@@ -321,7 +353,8 @@ and its totals.
|
|
|
321
353
|
|
|
322
354
|
**[The research series](./docs/blog/index.md)** ·
|
|
323
355
|
**[The harness](./harness/)** ·
|
|
324
|
-
**[Errata](./docs/blog/errata.md)**
|
|
356
|
+
**[Errata](./docs/blog/errata.md)** ·
|
|
357
|
+
**[Maturity milestones](./SPEC.md#91-maturity-milestones)**
|
|
325
358
|
|
|
326
359
|
## The architecture was earned by failures
|
|
327
360
|
|
|
@@ -362,7 +395,7 @@ via npm trusted publishing with SLSA provenance. Full rule:
|
|
|
362
395
|
|
|
363
396
|
```bash
|
|
364
397
|
npm install && npm run build # bundles the CLI to dist/cli/index.js
|
|
365
|
-
npm test #
|
|
398
|
+
npm test # 540+ tests — parser, detectors, engine, policy, renderers
|
|
366
399
|
npm run typecheck
|
|
367
400
|
```
|
|
368
401
|
|
package/dist/cli/index.js
CHANGED
|
@@ -33,6 +33,11 @@ function unquotePath(p) {
|
|
|
33
33
|
}
|
|
34
34
|
return Buffer.from(bytes).toString("utf8");
|
|
35
35
|
}
|
|
36
|
+
function pathToken(raw) {
|
|
37
|
+
if (raw.startsWith('"')) return raw;
|
|
38
|
+
const tab = raw.indexOf(" ");
|
|
39
|
+
return tab >= 0 ? raw.slice(0, tab) : raw;
|
|
40
|
+
}
|
|
36
41
|
function endOfQuoted(s) {
|
|
37
42
|
for (let i = 1; i < s.length; i++) {
|
|
38
43
|
if (s[i] === "\\") {
|
|
@@ -59,10 +64,17 @@ function headerPaths(line) {
|
|
|
59
64
|
first = rest.slice(0, q);
|
|
60
65
|
rest = rest.slice(q + 1);
|
|
61
66
|
} else {
|
|
62
|
-
const
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
67
|
+
const mid = (rest.length - 1) / 2;
|
|
68
|
+
const sym = Number.isInteger(mid) && rest[mid] === " " && rest.startsWith("a/") && rest.slice(mid + 1).startsWith("b/") && rest.slice(2, mid) === rest.slice(mid + 3);
|
|
69
|
+
if (sym) {
|
|
70
|
+
first = rest.slice(0, mid);
|
|
71
|
+
rest = rest.slice(mid + 1);
|
|
72
|
+
} else {
|
|
73
|
+
const m = rest.match(/^(.*) (b\/.*)$/);
|
|
74
|
+
if (!m) return [null, null];
|
|
75
|
+
first = m[1];
|
|
76
|
+
rest = m[2];
|
|
77
|
+
}
|
|
66
78
|
}
|
|
67
79
|
}
|
|
68
80
|
return [stripAB(first), stripAB(rest)];
|
|
@@ -126,8 +138,8 @@ function parseDiff(diff) {
|
|
|
126
138
|
renameTo = unquotePath(l.slice(8));
|
|
127
139
|
if (op === "modify") op = "rename";
|
|
128
140
|
} else if (l.startsWith("Binary files ") || l.startsWith("GIT binary patch")) binary = true;
|
|
129
|
-
else if (l.startsWith("--- ")) minusPath = l.slice(4);
|
|
130
|
-
else if (l.startsWith("+++ ")) plusPath = l.slice(4);
|
|
141
|
+
else if (l.startsWith("--- ")) minusPath = pathToken(l.slice(4));
|
|
142
|
+
else if (l.startsWith("+++ ")) plusPath = pathToken(l.slice(4));
|
|
131
143
|
i++;
|
|
132
144
|
}
|
|
133
145
|
let path;
|
|
@@ -196,7 +208,12 @@ function git(args, cwd) {
|
|
|
196
208
|
return execFileSync("git", args, {
|
|
197
209
|
cwd,
|
|
198
210
|
encoding: "utf8",
|
|
199
|
-
|
|
211
|
+
// 256 MiB, matching the hook adapter. Every diff below runs with `--text`, so a
|
|
212
|
+
// changed binary asset now costs its full size in patch output instead of one
|
|
213
|
+
// "Binary files differ" line; a legitimate golden-image update must fit. Past
|
|
214
|
+
// the cap git() throws and the view fails CLOSED (a truncated patch would parse
|
|
215
|
+
// as a smaller change than the one landing), which is the documented stance.
|
|
216
|
+
maxBuffer: 256 * 1024 * 1024,
|
|
200
217
|
stdio: ["ignore", "pipe", "pipe"]
|
|
201
218
|
});
|
|
202
219
|
} catch (e) {
|
|
@@ -220,7 +237,7 @@ function fromDisk(path, cwd) {
|
|
|
220
237
|
}
|
|
221
238
|
}
|
|
222
239
|
function enrich(c, beforeReader, afterReader) {
|
|
223
|
-
if (c.kind !== "file"
|
|
240
|
+
if (c.kind !== "file") return c;
|
|
224
241
|
const before = c.op !== "add" ? beforeReader(c.oldPath ?? c.path) : null;
|
|
225
242
|
const after = c.op !== "delete" ? afterReader(c.path) : null;
|
|
226
243
|
return { ...c, before, after };
|
|
@@ -261,8 +278,9 @@ function headSha(cwd) {
|
|
|
261
278
|
return null;
|
|
262
279
|
}
|
|
263
280
|
}
|
|
281
|
+
var DIFF_ARGS = ["diff", "--no-color", "--text", "-M"];
|
|
264
282
|
function diffSince(rev, opts = {}) {
|
|
265
|
-
const raw = git([
|
|
283
|
+
const raw = git([...DIFF_ARGS, rev], opts.cwd);
|
|
266
284
|
return parseDiff(raw).map(
|
|
267
285
|
(c) => enrich(c, (p) => blobAt(rev, p, opts.cwd), (p) => fromDisk(p, opts.cwd))
|
|
268
286
|
);
|
|
@@ -270,20 +288,20 @@ function diffSince(rev, opts = {}) {
|
|
|
270
288
|
function diffRange(base, head, opts = {}) {
|
|
271
289
|
assertRev(base);
|
|
272
290
|
assertRev(head);
|
|
273
|
-
const raw = git([
|
|
291
|
+
const raw = git([...DIFF_ARGS, `${base}...${head}`], opts.cwd);
|
|
274
292
|
const mergeBase = mergeBaseOf(base, head, opts);
|
|
275
293
|
return parseDiff(raw).map(
|
|
276
294
|
(c) => enrich(c, (p) => blobAt(mergeBase, p, opts.cwd), (p) => blobAt(head, p, opts.cwd))
|
|
277
295
|
);
|
|
278
296
|
}
|
|
279
297
|
function diffStaged(opts = {}) {
|
|
280
|
-
const raw = git([
|
|
298
|
+
const raw = git([...DIFF_ARGS, "--cached"], opts.cwd);
|
|
281
299
|
return parseDiff(raw).map(
|
|
282
300
|
(c) => enrich(c, (p) => blobAt("HEAD", p, opts.cwd), (p) => blobAt("", p, opts.cwd))
|
|
283
301
|
);
|
|
284
302
|
}
|
|
285
303
|
function diffWorktree(opts = {}) {
|
|
286
|
-
const raw = git([
|
|
304
|
+
const raw = git([...DIFF_ARGS, "HEAD"], opts.cwd);
|
|
287
305
|
return parseDiff(raw).map(
|
|
288
306
|
(c) => enrich(c, (p) => blobAt("HEAD", p, opts.cwd), (p) => fromDisk(p, opts.cwd))
|
|
289
307
|
);
|
|
@@ -1931,15 +1949,23 @@ function applyLocalSignoffs(findings, cwd, policy, now = Date.now()) {
|
|
|
1931
1949
|
}
|
|
1932
1950
|
return { findings: remaining, cleared };
|
|
1933
1951
|
}
|
|
1934
|
-
function
|
|
1935
|
-
const
|
|
1936
|
-
|
|
1952
|
+
function oobToken(want, oob, head) {
|
|
1953
|
+
for (const raw of oob) {
|
|
1954
|
+
const t = raw.trim();
|
|
1955
|
+
if (!t) continue;
|
|
1937
1956
|
const at = t.lastIndexOf("@");
|
|
1938
|
-
if (at === -1)
|
|
1957
|
+
if (at === -1) {
|
|
1958
|
+
if (!head && t === want) return t;
|
|
1959
|
+
continue;
|
|
1960
|
+
}
|
|
1939
1961
|
const [rule, sha2] = [t.slice(0, at), t.slice(at + 1)];
|
|
1940
|
-
if (rule !== want)
|
|
1941
|
-
|
|
1942
|
-
}
|
|
1962
|
+
if (rule !== want) continue;
|
|
1963
|
+
if (!head || sha2.length >= 7 && head.startsWith(sha2)) return t;
|
|
1964
|
+
}
|
|
1965
|
+
return null;
|
|
1966
|
+
}
|
|
1967
|
+
function applyOobSignoffs(findings, oob, head) {
|
|
1968
|
+
const matches = (want) => oobToken(want, oob, head) !== null;
|
|
1943
1969
|
const cleared = [];
|
|
1944
1970
|
const remaining = [];
|
|
1945
1971
|
for (const f of findings) {
|
|
@@ -2294,7 +2320,7 @@ function synthFileChange(displayPath, before, after) {
|
|
|
2294
2320
|
writeFileSync(a, before ?? "");
|
|
2295
2321
|
writeFileSync(b, after ?? "");
|
|
2296
2322
|
try {
|
|
2297
|
-
raw = execFileSync2("git", ["diff", "--no-index", "--no-color", a, b], {
|
|
2323
|
+
raw = execFileSync2("git", ["diff", "--no-index", "--no-color", "--text", a, b], {
|
|
2298
2324
|
encoding: "utf8",
|
|
2299
2325
|
maxBuffer: 256 * 1024 * 1024
|
|
2300
2326
|
});
|
|
@@ -3025,8 +3051,9 @@ var WORKFLOW_CONTENT = `name: tamperward
|
|
|
3025
3051
|
|
|
3026
3052
|
# The CI authority for main: the same engine as the agent hook and pre-commit, run over
|
|
3027
3053
|
# the PR's commit range. A block fails the check and clears ONLY via the out-of-band
|
|
3028
|
-
# label \`tamperward:allow:<rule>\` applied by someone with
|
|
3029
|
-
# the PR itself can commit.
|
|
3054
|
+
# label \`tamperward:allow:<rule>@<head-sha>\` applied by someone with triage access or higher \u2014
|
|
3055
|
+
# never a file the PR itself can commit. The verify step reads the same labels:
|
|
3056
|
+
# \`tamperward:allow:verify@<head-sha>\` accepts a masked failure a reviewer has judged.
|
|
3030
3057
|
#
|
|
3031
3058
|
# labeled/unlabeled re-run the gate because the sign-off is read from the EVENT payload:
|
|
3032
3059
|
# a label applied after a failure could otherwise never take effect, and REVOKING a
|
|
@@ -3082,6 +3109,14 @@ jobs:
|
|
|
3082
3109
|
# Requires a \`verify:\` block in .tamperward.yml naming the suite command;
|
|
3083
3110
|
# without one this step fails closed (exit 2) rather than passing quietly.
|
|
3084
3111
|
- name: Tamperward verify (pristine re-execution)
|
|
3112
|
+
env:
|
|
3113
|
+
# The same channel as the gate. \`tamperward:allow:verify@<head-sha>\`
|
|
3114
|
+
# accepts a MASKED_FAILURE \u2014 a reviewer has read the intentional test
|
|
3115
|
+
# change and agrees the original suite no longer applies. It clears
|
|
3116
|
+
# nothing else: a red visible suite, or a run that could not verify,
|
|
3117
|
+
# stays red whatever the labels say.
|
|
3118
|
+
TAMPERWARD_OOB_SIGNOFF: \${{ steps.oob.outputs.rules }}
|
|
3119
|
+
TAMPERWARD_OOB_HEAD: \${{ github.event.pull_request.head.sha }}
|
|
3085
3120
|
run: npx --yes tamperward@${TW_VERSION} verify --require-ancestor --base "\${{ github.event.pull_request.base.sha }}"
|
|
3086
3121
|
`;
|
|
3087
3122
|
var WORKFLOW_MARK = "# tamperward:generated";
|
|
@@ -3768,6 +3803,8 @@ function runVerify(opts) {
|
|
|
3768
3803
|
verdict2 = "SUITE_RED";
|
|
3769
3804
|
code = 1;
|
|
3770
3805
|
}
|
|
3806
|
+
const signedOff = verdict2 === "MASKED_FAILURE" ? oobToken("verify", oobFromEnv(), oobHeadFromEnv()) : null;
|
|
3807
|
+
if (signedOff) code = 0;
|
|
3771
3808
|
if (opts.json) {
|
|
3772
3809
|
out2(
|
|
3773
3810
|
JSON.stringify({
|
|
@@ -3779,6 +3816,7 @@ function runVerify(opts) {
|
|
|
3779
3816
|
pristine: { exit: pristine.exit, secs: pristine.secs },
|
|
3780
3817
|
protected_restored: restored.length,
|
|
3781
3818
|
added_protected_removed: removedAdded,
|
|
3819
|
+
...signedOff ? { oob_signoff: signedOff } : {},
|
|
3782
3820
|
...opts.keep ? { visible_dir: visDir, pristine_dir: priDir } : {}
|
|
3783
3821
|
})
|
|
3784
3822
|
);
|
|
@@ -3790,6 +3828,10 @@ function runVerify(opts) {
|
|
|
3790
3828
|
BUDGET_EXCEEDED: `budget exceeded (${budget}s): could not verify \u2014 failing closed, not open.`
|
|
3791
3829
|
};
|
|
3792
3830
|
out2(`tamperward verify \u2014 ${lines[verdict2]}`);
|
|
3831
|
+
if (signedOff)
|
|
3832
|
+
out2(
|
|
3833
|
+
`masked failure cleared by out-of-band approval (tamperward:allow:${signedOff}): a reviewer accepted that the original suite no longer applies to this change. Exit 0.`
|
|
3834
|
+
);
|
|
3793
3835
|
if (removedAdded > 0)
|
|
3794
3836
|
out2(
|
|
3795
3837
|
`(${removedAdded} protected file(s) added since ${base.slice(0, 10)} were removed from the pristine run: the pristine tree carries exactly the base's protected surface.)`
|
|
@@ -4077,20 +4119,24 @@ Formats:
|
|
|
4077
4119
|
observe supported transient effects
|
|
4078
4120
|
tamperward verify [--base R] [--cmd C] pristine-suite re-execution: run the suite
|
|
4079
4121
|
[--budget S] [--json] [--keep] as-is AND with protected files restored
|
|
4080
|
-
|
|
4122
|
+
[--require-ancestor] [--cwd D] from the trusted base; a visible-green /
|
|
4081
4123
|
pristine-red result is a MASKED FAILURE
|
|
4082
|
-
(exit 1
|
|
4124
|
+
(exit 1, or 0 under an out-of-band
|
|
4125
|
+
verify@<head-sha> approval); cannot-verify
|
|
4126
|
+
fails closed (2)
|
|
4083
4127
|
tamperward run [opts] -- <agent cmd...> enforcement envelope: record the trusted
|
|
4084
4128
|
[--base R] [--cmd C] base, run the agent, treat its exit as
|
|
4085
4129
|
[--budget S] [--allow-dirty] untrusted, then re-adjudicate the tree it
|
|
4086
|
-
|
|
4130
|
+
[--settle S] [--allow-dep-drift] left (policy over base...HEAD and the
|
|
4131
|
+
[--cwd D]
|
|
4087
4132
|
worktree, plus verify). Exit: agent's code
|
|
4088
4133
|
when clean; 1 on any blocking finding or
|
|
4089
4134
|
masked failure; 2 when it cannot
|
|
4090
4135
|
adjudicate (fails closed)
|
|
4091
4136
|
tamperward allow <rule> --reason "..." record a human sign-off (local audit ledger)
|
|
4137
|
+
[--file F] [--cwd D]
|
|
4092
4138
|
tamperward init [--dry-run] wire the policy file plus every
|
|
4093
|
-
[--force-workflow]
|
|
4139
|
+
[--force-workflow] [--cwd D] enforcement point: Claude Code hooks,
|
|
4094
4140
|
pre-commit, CI. Idempotent; never
|
|
4095
4141
|
overwrites your files. Re-running it
|
|
4096
4142
|
MIGRATES a CI workflow it generated and
|
|
@@ -4100,8 +4146,12 @@ Formats:
|
|
|
4100
4146
|
--force-workflow replaces a workflow it
|
|
4101
4147
|
did not write, or one you have edited.
|
|
4102
4148
|
|
|
4103
|
-
Exit
|
|
4104
|
-
|
|
4149
|
+
Exit codes: 0 clean \xB7 1 a blocking finding (check), MASKED_FAILURE or SUITE_RED
|
|
4150
|
+
(verify), any blocking finding or masked failure (run) \xB7 2 cannot
|
|
4151
|
+
evaluate, failing closed: policy error, bad --diff range, no view,
|
|
4152
|
+
unresolvable base, --require-ancestor refused, budget exceeded.
|
|
4153
|
+
hook/sweep \u2192 0 with a deny emitted as JSON on stdout (exit 2 would make
|
|
4154
|
+
Claude Code ignore the JSON); 2 only for an unsupported agent name.
|
|
4105
4155
|
`);
|
|
4106
4156
|
}
|
|
4107
4157
|
function main(argv2) {
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "tamperward",
|
|
3
|
-
"version": "2.
|
|
3
|
+
"version": "2.1.1",
|
|
4
4
|
"description": "The deterministic agent-integrity gate. One ruleset, evaluated on the actual diff/commands as a verdict, enforced everywhere a change can be made.",
|
|
5
5
|
"license": "Apache-2.0",
|
|
6
6
|
"author": "hexrift",
|