tamperward 2.21.0 → 2.23.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +20 -2
- package/dist/cli/index.js +7209 -4928
- package/package.json +1 -1
- package/schemas/research-v1.schema.json +679 -0
package/README.md
CHANGED
|
@@ -259,6 +259,7 @@ platforms.
|
|
|
259
259
|
| `check` / policy evaluation | Supported | Supported | Supported |
|
|
260
260
|
| Claude hook / Stop adapter | Supported where Claude Code command hooks are available | Same | Same |
|
|
261
261
|
| `watch` / observer telemetry | Supported; backend health is reported | Supported/degraded according to `fs.watch` health | Supported/degraded according to `fs.watch` health |
|
|
262
|
+
| opt-in `hook-service` | Supported (per-user `0600` unix socket) | Supported (per-user `0600` unix socket) | **Unsupported; `start` refuses, hooks run in-process** |
|
|
262
263
|
| checkpointed-local `verify` | Supported via `/bin/sh` | Supported via `/bin/sh` | **Unsupported; fails before candidate execution** |
|
|
263
264
|
| isolated-container `verify` | Supported when Docker authority preflight passes | Not claimed beyond Docker preflight | Not claimed beyond Docker preflight |
|
|
264
265
|
| advisory `trace-verify` | **Supported with `strace` + `tar`** | **Unsupported; reports no parity** | **Unsupported; reports no parity** |
|
|
@@ -522,14 +523,15 @@ from trusted CI, or after an externally isolated agent hands off the frozen cand
|
|
|
522
523
|
### Machine-readable verdict API
|
|
523
524
|
|
|
524
525
|
From **2.19.0**, the public JSON verdict surfaces are versioned independently of the
|
|
525
|
-
npm package version. `check --json`, `verify --json`, `run --json`,
|
|
526
|
-
`doctor --json` include top-level `"schema_version": 1`. TamperWard publishes the
|
|
526
|
+
npm package version. `check --json`, `verify --json`, `run --json`,
|
|
527
|
+
`doctor --json`, `research run --json` and `research summarize` include top-level `"schema_version": 1`. TamperWard publishes the
|
|
527
528
|
corresponding JSON Schema Draft 2020-12 documents in the npm package and repository:
|
|
528
529
|
|
|
529
530
|
- [`schemas/check-v1.schema.json`](./schemas/check-v1.schema.json)
|
|
530
531
|
- [`schemas/verify-v1.schema.json`](./schemas/verify-v1.schema.json)
|
|
531
532
|
- [`schemas/run-v1.schema.json`](./schemas/run-v1.schema.json)
|
|
532
533
|
- [`schemas/doctor-v1.schema.json`](./schemas/doctor-v1.schema.json)
|
|
534
|
+
- [`schemas/research-v1.schema.json`](./schemas/research-v1.schema.json) — from **2.23.0**, the `pair` records `research run` writes (and prints with `--json`) and the `summary` document `research summarize` prints
|
|
533
535
|
|
|
534
536
|
Schema major **1** is deliberately additive: consumers should ignore fields they do not
|
|
535
537
|
understand. Adding new evidence/diagnostic fields does not require a schema bump.
|
|
@@ -580,10 +582,13 @@ option can never be reinterpreted as the agent command.
|
|
|
580
582
|
| `trace-verify` | Linux-only advisory discovery: `--base <rev>` (default `HEAD`) · `--cmd <suite command>` · `--budget <seconds>` · `--runs <positive integer>` (default 2) · `--json` · `--cwd <dir>` |
|
|
581
583
|
| `doctor` | `--base <rev>` (trusted policy revision) · `--workflow <path>` · `--cwd <dir>` · `--json` · `--github` · `--repo <owner/repo>` · `--branch <name>` — read-only installation/authority posture plus CI verifier outer-time validation |
|
|
582
584
|
| `run` | `--base <rev>` · `--cmd <suite command>` · `--budget <seconds>` (per verifier suite) · `--agent-budget <seconds>` (optional wrapped-agent wall clock) · `--json` (one versioned final envelope document) · `--observe-transients` (start a session-scoped transient observer) · `--allow-dirty` · `--settle <seconds>` (wait before the final quiescence check) · `--allow-dep-drift` · `--cwd <dir>` · then `-- <agent command...>` |
|
|
585
|
+
| `research run` | `--manifest <file>` · `--out <dir>` · `--adapter claude-code\|command` (all three required) · `--pairs <n>` · `--model <id>` · `--agent-budget <seconds>` · `--json` (one pair record per line) · then `-- <agent command...>` for the `command` adapter, with `{prompt}` `{task}` `{cwd}` `{base}` `{arm}` `{model}` substituted — see [the research guide](./docs/guide/research.md) |
|
|
586
|
+
| `research summarize` | `--ledger <dir>` (required) — one aggregate document, four separated readouts, no composite score |
|
|
583
587
|
| `allow` | `<rule>` · `--file <path>` · `--reason "<why>"` (required) · `--cwd <dir>` |
|
|
584
588
|
| `init` | `--cwd <dir>` · `--dry-run` · `--force-workflow` |
|
|
585
589
|
| `onboard` | `--cwd <dir>` · `--base <rev>` · `--repo <owner/repo>` · `--branch <name>` · `--skip-demo` / `--demo` (mutually exclusive) · `--no-github` · `--yes` (scripted: no prompts; the demo runs only with `--demo`) · `--verify-command "<suite command>"` (the only way a scripted run configures `verify.command`) |
|
|
586
590
|
| `watch` | `--dir <dir>` · `--log <file>` — a daemon; it runs until signalled |
|
|
591
|
+
| `hook-service` | `start [--dir <repo>]` (foreground; runs until signalled) · `stop` · `status` — the opt-in persistent hook service (2.22.0): one warm process per user, bound to one repository that evaluates `hook`/`sweep` payloads over a private `0600` unix socket. Hooks consult it only under `TAMPERWARD_HOOK_SERVICE=1`. Before handoff, unavailable/refusing service paths fall back to the same in-process verdict; after handoff, ambiguous transport failure fails closed rather than starting a concurrent second evaluation. Not available on Windows |
|
|
587
592
|
| `hook claude` / `sweep claude` | none — the Claude Code payload arrives on stdin |
|
|
588
593
|
|
|
589
594
|
**Exit codes** — part of the public surface:
|
|
@@ -595,7 +600,9 @@ option can never be reinterpreted as the agent command.
|
|
|
595
600
|
| `trace-verify` | all requested known-good traces exited 0; advisory report emitted | one or more traced verifier runs were non-zero/incomplete; report still emitted | unsupported platform, missing tracer/materialiser, bad base/policy/options, or tracing failure | — |
|
|
596
601
|
| `doctor` | configured verify job(s) have sufficient static outer time for the trusted policy | — | missing/invalid workflow, no verify job, missing/malformed/insufficient timeout, or trusted policy cannot be loaded | — |
|
|
597
602
|
| `run` | enforcement clean and the agent exited 0 — another non-zero agent exit is passed through unchanged | any blocking finding or masked failure, including a non-quiescent process after timeout | cannot adjudicate: dirty start, policy error, verify cannot run | `AGENT_TIMEOUT`: `--agent-budget` expired and post-timeout enforcement was clean |
|
|
603
|
+
| `research run` / `research summarize` | every requested pair recorded (or already was); summary printed | — | cannot start or set a trajectory up: bad manifest, unknown adapter, root or unsupported platform (doctor's own `platform` check), unclonable repository, or invalid/mixed ledger evidence — the agent's own exit is data in the record, never the research exit | — |
|
|
598
604
|
| `hook claude` / `sweep claude` | always — a deny is JSON on stdout at exit 0, never exit 2 | — | only for an unsupported agent name | — |
|
|
605
|
+
| `hook-service` | started, stopped (or nothing to stop), or status printed | — | unsupported platform, a runtime directory another uid owns, or a service already listening | — |
|
|
599
606
|
| `allow` | sign-off recorded | — | no rule or `--reason`, not a git repo, or no current blocking finding to sign off | — |
|
|
600
607
|
| `init` | wired, or already wired | — | an item needs attention | — |
|
|
601
608
|
| `onboard` | posture `READY` or `READY WITH WARNINGS` (from `doctor`) | posture `BROKEN` or `INCOMPLETE`, including a declined write or an unconfigured verifier | refused (not a git repository, non-interactive stdin without `--yes`, a dirty tree the operator would not continue on) or aborted at a prompt | — |
|
|
@@ -609,6 +616,8 @@ option can never be reinterpreted as the agent command.
|
|
|
609
616
|
| `TAMPERWARD_OOB_HEAD` | the CI workflow (`github.event.pull_request.head.sha`) | the head SHA under adjudication; once set, an approval clears anything only if it names that commit (`@<sha>`, at least 7 characters), so a new push re-blocks |
|
|
610
617
|
| `TAMPERWARD_DENYLOG` | a harness or operator | a file to which `hook claude` and `sweep claude` append the rule ids of every deny, one line per verdict, best effort |
|
|
611
618
|
| `TAMPERWARD_FSEVENTS` | operator or harness | overrides the `tamperward watch` event-log path (default `.git/tamperward/fsevents.jsonl`); the Stop sweep reads the same variable |
|
|
619
|
+
| `TAMPERWARD_HOOK_SERVICE` | the operator, in Claude Code's environment (`=1`) | lets the hooks hand their payload to a running `tamperward hook-service`; off by default. Pre-handoff refusal falls back to in-process evaluation; post-handoff ambiguity fails closed so two evaluations never race one session |
|
|
620
|
+
| `TAMPERWARD_HOOK_SERVICE_DIR` | the operator or tests | overrides the service's runtime directory (default `$XDG_RUNTIME_DIR/tamperward-hook`, else `<tmpdir>/tamperward-hook-<uid>`); it must be the hook's own uid at `0700`, the socket `0600` |
|
|
612
621
|
| `TAMPERWARD_WATCH_NO_RECURSIVE` | CI and tests (`=1`) | forces `tamperward watch` onto its per-directory fallback instead of recursive `fs.watch`, so the fallback is exercised on every platform |
|
|
613
622
|
| `TAMPERWARD_TRANSIENT` | a harness that owns restore semantics (`=block`) | raises `transient-protected-mutation` from warn to block; it can never lower a severity |
|
|
614
623
|
| `NO_COLOR` / `FORCE_COLOR` | the user's shell | any non-empty `NO_COLOR` disables colour in the text renderer; a non-empty, non-`0` `FORCE_COLOR` enables it |
|
|
@@ -677,6 +686,15 @@ predictions and corrections remain in the public record rather than being
|
|
|
677
686
|
removed after the result is known; the series and its errata carry the ledger
|
|
678
687
|
and its totals.
|
|
679
688
|
|
|
689
|
+
Since **2.23.0** the paired ungated/gated evaluation is a supported command rather
|
|
690
|
+
than a hand-assembled harness run: `tamperward research run` takes a task manifest
|
|
691
|
+
and an agent runtime (Claude Code, or any command through the `AgentAdapter`
|
|
692
|
+
contract), pins one source commit per task, adjudicates both arms with the same
|
|
693
|
+
`verify` + `check` primitives, records TamperWard's own verdict separately from that
|
|
694
|
+
outcome, and `research summarize` reports model behaviour, independent outcome,
|
|
695
|
+
TamperWard hits/misses and paired counts with no composite score — a result in which
|
|
696
|
+
TamperWard loses is as plain as one in which it wins. **[Guide](./docs/guide/research.md)**.
|
|
697
|
+
|
|
680
698
|
**[The research series](./docs/blog/index.md)** ·
|
|
681
699
|
**[The harness](./harness/)** ·
|
|
682
700
|
**[Errata](./docs/blog/errata.md)** ·
|