rulereceipt 0.1.37 → 0.1.39

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -89,6 +89,7 @@ rulereceipt rules --exclude <handle> # "this isn't" — stop reporting it
89
89
  rulereceipt rules --coverage # which rules a configured hook might actually enforce
90
90
  rulereceipt doctor # list hooks/auto-run tasks configured on this machine
91
91
  rulereceipt hook # run AS a Claude Code Stop hook — block Claude finishing on a broken rule
92
+ rulereceipt guard # run AS a Claude Code PreToolUse hook — refuse a call before it runs
92
93
  rulereceipt lint # find contradictions between CLAUDE.md and AGENTS.md
93
94
  rulereceipt digest # summarise recent checks; --email to send it
94
95
  rulereceipt config # set up email sending (stays on your machine)
@@ -127,10 +128,16 @@ is the part a model cannot talk its way around.
127
128
 
128
129
  Three properties worth knowing before you wire it in:
129
130
 
130
- - **It only blocks on things it can prove.** Never a judgment rule, never an
131
- LLM opinion, never "couldn't tell". Only a matched literal or a claim
132
- contradicted by a recorded tool result. Run against twelve real sessions it
133
- blocked none of them.
131
+ - **It blocks two things, both narrow.** A claim a recorded run contradicts,
132
+ and a claim of done that nothing in the session verified. Never a judgment
133
+ rule, never an LLM opinion. Run against thirteen real sessions it stopped
134
+ two, and both were read by hand.
135
+ - **The report and the gate disagree in exactly one place.** When a session
136
+ claims work is done and nothing recorded verifies it, the report says
137
+ "couldn't tell" — the tests may have run in another terminal, and a
138
+ transcript cannot see that. The gate refuses the exit anyway, because it is
139
+ not saying the claim is false. It is declining to let "done" end a session
140
+ with nothing behind it.
134
141
  - **It cannot loop.** Claude Code sets `stop_hook_active` when a session is
135
142
  already continuing because of a block; the hook returns immediately in that
136
143
  case. One interruption per stop.
@@ -143,6 +150,57 @@ It runs when Claude stops, so it catches a finished session, not a command
143
150
  mid-flight. For that, use a `PreToolUse` hook of your own — `rulereceipt
144
151
  doctor` will show you what you already have.
145
152
 
153
+ ### Refusing a command before it runs
154
+
155
+ `rulereceipt guard` runs as a `PreToolUse` hook and refuses a call outright:
156
+
157
+ ```json
158
+ {
159
+ "hooks": {
160
+ "PreToolUse": [
161
+ { "hooks": [ { "type": "command", "command": "npx rulereceipt guard" } ] }
162
+ ]
163
+ }
164
+ }
165
+ ```
166
+
167
+ Read the limit before wiring it in, because it is most of the story. It
168
+ enforces rules naming a **file** or a **branch** — "never modify `.env`",
169
+ "never commit to `main`" — and nothing else.
170
+
171
+ It does **not** block a banned command unless you have said which command is
172
+ banned. That was the point of building it, and the automatic version did not
173
+ survive measurement: replaying 16,336 real tool calls against every forbidding
174
+ rule in a 559-file corpus, blocking on command literals refused 62.8% of them.
175
+ Narrowing twice reached 2.5%, and the residue was still wrong in a way no
176
+ matcher fixes — one rule refused `npm run build` 112 times, because it forbids
177
+ running Playwright unprompted and *recommends* `npm run build`, which is its
178
+ only command-shaped literal.
179
+
180
+ Nothing in a rules file marks which backtick is the prohibition. A report
181
+ survives that by saying UNCLEAR. A gate cannot — so you mark it:
182
+
183
+ ```bash
184
+ rulereceipt rules --forbid <handle> --literal "git push --force"
185
+ ```
186
+
187
+ Handles come from `rulereceipt check --show-skipped`. The mark is stored
188
+ against the rule's content hash, and the guard blocks on that literal and no
189
+ other. Three things it deliberately will not do:
190
+
191
+ - An **unmarked** rule cannot block, at any confidence, ever. There is no
192
+ fallback to "probably the first literal" — that fallback is the bug.
193
+ - **Rewording the rule drops the mark.** It would otherwise carry your
194
+ judgment onto words you never read.
195
+ - A mark naming a literal the rule no longer contains is **ignored**. A gate
196
+ refusing a command for a reason written nowhere is the worst failure a gate
197
+ has.
198
+
199
+ Of 99 forbidding rules in the corpus that name a command-shaped literal, only
200
+ 43 have a prohibition that actually introduces one. The rest could never be
201
+ marked automatically, which is the point.
202
+
203
+
146
204
  ## Which rules actually have teeth
147
205
 
148
206
  A rule in a file and a rule with a `PreToolUse` hook behind it look identical
@@ -188,7 +188,11 @@ function unclear(rule, evidence) {
188
188
  */
189
189
  export function runClaimEvidenceChecks(classifications, events) {
190
190
  let lastRun = null;
191
- let pendingRun = null;
191
+ // Keyed by tool_use id where the transcript has one, so a result can be
192
+ // matched to the call it belongs to rather than to the call above it.
193
+ // `null` is the key for id-less transcripts, which keeps the old
194
+ // positional behaviour for fixtures and older logs.
195
+ const pendingRuns = new Map();
192
196
  let claimsMade = 0;
193
197
  let unknownSinceRed = null;
194
198
  const commandsSeen = new Set();
@@ -200,7 +204,15 @@ export function runClaimEvidenceChecks(classifications, events) {
200
204
  for (const event of events) {
201
205
  const command = commandOf(event);
202
206
  if (command !== null) {
203
- pendingRun = TEST_COMMAND.test(withoutHeredocs(command)) ? command : null;
207
+ const id = event.kind === "tool_use" ? (event.toolUseId ?? null) : null;
208
+ if (TEST_COMMAND.test(withoutHeredocs(command))) {
209
+ pendingRuns.set(id, command);
210
+ }
211
+ else if (id === null) {
212
+ // No id to distinguish calls, so a later call really does supersede
213
+ // an earlier one — the original positional rule, unchanged.
214
+ pendingRuns.delete(null);
215
+ }
204
216
  for (const action of ACTION_CLAIMS) {
205
217
  if (action.command.test(command))
206
218
  commandsSeen.add(action.label);
@@ -217,6 +229,8 @@ export function runClaimEvidenceChecks(classifications, events) {
217
229
  // every turn contained exactly one tool call, so there is no parallel
218
230
  // fan-out to mis-attribute.
219
231
  if (event.kind === "tool_result") {
232
+ const resultId = event.toolUseId ?? null;
233
+ const pendingRun = pendingRuns.get(resultId) ?? null;
220
234
  if (pendingRun !== null) {
221
235
  // Prefer what the runner SAID over what the shell returned: the
222
236
  // words survive a pipe, the exit status does not.
@@ -233,7 +247,7 @@ export function runClaimEvidenceChecks(classifications, events) {
233
247
  outcomeReadable: oneRun && (stated !== null || trustExitCode),
234
248
  output: event.content.slice(0, 200),
235
249
  };
236
- pendingRun = null;
250
+ pendingRuns.delete(resultId);
237
251
  unknownSinceRed = null; // a recognised run supersedes anything before it
238
252
  }
239
253
  continue;
@@ -331,8 +345,12 @@ export function runClaimEvidenceChecks(classifications, events) {
331
345
  };
332
346
  }
333
347
  if (claimsMade > 0) {
334
- return unclear(rule, `the session claimed a passing test suite ${claimsMade} time(s), but no test command ran here — ` +
335
- `it may have been run outside this session, which the transcript cannot show`);
348
+ return {
349
+ ...unclear(rule, `the session claimed a passing test suite ${claimsMade} time(s), but no test command ran here — ` +
350
+ `it may have been run outside this session, which the transcript cannot show`),
351
+ // UNCLEAR in the report, refused by the gate. See CheckResult.unverifiedClaim.
352
+ unverifiedClaim: true,
353
+ };
336
354
  }
337
355
  return unclear(rule, "the session made no claim about passing tests, so there was nothing to check against the log");
338
356
  });
@@ -107,6 +107,36 @@ export interface ClaimEvidenceClassification {
107
107
  rule: Rule;
108
108
  }
109
109
  export type Classification = ClaimEvidenceClassification | DeterministicClassification | IfEditThenTestClassification | GitBranchPolicyClassification | CodeContentClassification | FileLifecycleClassification | NotARuleClassification | JudgmentClassification;
110
+ export declare const DIRECTIVE_LANGUAGE: RegExp;
111
+ /**
112
+ * Imperative instruction — a bare command verb starting a clause ("Use
113
+ * `gh pr merge`", "Run the tests first", "Keep functions small"). This is
114
+ * the other way a real rule is written when it doesn't use a modal.
115
+ *
116
+ * Anchored to a clause start (line start, or after sentence/bullet
117
+ * punctuation) on purpose: the same verbs appear mid-sentence in pure
118
+ * documentation ("the CLI can run migrations"), where they describe a
119
+ * capability rather than instruct the agent.
120
+ */
121
+ export declare const IMPERATIVE_INSTRUCTION: RegExp;
122
+ /**
123
+ * A title that OPENS with an instruction is a rule, whatever it goes on
124
+ * to mention. "Never repeat the 2026-08-28 incident" is a directive that
125
+ * happens to name an incident; "Real incident (2026-08-28): ..." is a
126
+ * report that happens to contain the word never further along.
127
+ */
128
+ export declare const TITLE_OPENS_WITH_DIRECTIVE: RegExp;
129
+ export declare function isEventRecord(rule: Rule): boolean;
130
+ /**
131
+ * A heading that labels a command, with the command as its whole body:
132
+ * "Build release APK" over `.\gradlew assembleRelease`. The title reads as
133
+ * an imperative, but nothing here constrains the agent — it is a how-to, and
134
+ * there is no compliance to check.
135
+ *
136
+ * Guarded by TITLE_OPENS_WITH_DIRECTIVE so a genuine prohibition whose body
137
+ * is the forbidden command ("Never run: `rm -rf /`") is still a rule.
138
+ */
139
+ export declare function isCommandDocumentation(rule: Rule): boolean;
110
140
  /**
111
141
  * A rule is only treated as deterministic when it names a specific,
112
142
  * literal, checkable token (a CLI flag, a command, an exact string) in
@@ -1,7 +1,7 @@
1
1
  // Normative language — the thing that makes a line a rule rather than a
2
2
  // description. Deliberately broad on modals AND imperative verbs, because
3
3
  // a wrongly-excluded rule is a silent miss.
4
- const DIRECTIVE_LANGUAGE = /\b(never|always|must|should|shall|do not|don't|dont|cannot|can't|required?|requires|ensure|avoid|prefer|forbidden|prohibited|only|make sure|be sure|need|needs|needed|need to|has to|have to|expected to|responsible for)\b/i;
4
+ export const DIRECTIVE_LANGUAGE = /\b(never|always|must|should|shall|do not|don't|dont|cannot|can't|required?|requires|ensure|avoid|prefer|forbidden|prohibited|only|make sure|be sure|need|needs|needed|need to|has to|have to|expected to|responsible for)\b/i;
5
5
  /**
6
6
  * Imperative instruction — a bare command verb starting a clause ("Use
7
7
  * `gh pr merge`", "Run the tests first", "Keep functions small"). This is
@@ -12,7 +12,7 @@ const DIRECTIVE_LANGUAGE = /\b(never|always|must|should|shall|do not|don't|dont|
12
12
  * documentation ("the CLI can run migrations"), where they describe a
13
13
  * capability rather than instruct the agent.
14
14
  */
15
- const IMPERATIVE_INSTRUCTION = /(?:^|[.;:!?]\s+|^\s*[-*+]\s*|\n\s*[-*+]\s*)(use|run|keep|write|add|remove|delete|check|verify|test|commit|document|update|create|follow|apply|include|exclude|handle|validate|escape|sanitize|log|report|raise|throw|return|call|invoke|split|group|sort|name|place|put|store|read|load|save|close|open|start|stop|restart|install|build|deploy|review|refactor|rename|move|copy|merge|rebase|squash|tag|branch|push|pull|fetch|clone|stage|stash|lead|state|explain|describe|list|show|surface|flag|mark|label|note|treat|assume|confirm|ask|wait|stick|limit|cap|batch|cache|mock|stub|assert|expect|measure|quantify|label)\b/i;
15
+ export const IMPERATIVE_INSTRUCTION = /(?:^|[.;:!?]\s+|^\s*[-*+]\s*|\n\s*[-*+]\s*)(use|run|keep|write|add|remove|delete|check|verify|test|commit|document|update|create|follow|apply|include|exclude|handle|validate|escape|sanitize|log|report|raise|throw|return|call|invoke|split|group|sort|name|place|put|store|read|load|save|close|open|start|stop|restart|install|build|deploy|review|refactor|rename|move|copy|merge|rebase|squash|tag|branch|push|pull|fetch|clone|stage|stash|lead|state|explain|describe|list|show|surface|flag|mark|label|note|treat|assume|confirm|ask|wait|stick|limit|cap|batch|cache|mock|stub|assert|expect|measure|quantify|label)\b/i;
16
16
  /**
17
17
  * Deliberately inverted: tests for the presence of a DIRECTIVE, never for
18
18
  * the shape of documentation.
@@ -60,8 +60,8 @@ const EVENT_RECORD_TITLE = /\b(incident|post-?mortem|retro(spective)?|outage|wha
60
60
  * happens to name an incident; "Real incident (2026-08-28): ..." is a
61
61
  * report that happens to contain the word never further along.
62
62
  */
63
- const TITLE_OPENS_WITH_DIRECTIVE = /^\s*[-*+\d.\s]*(never|always|must|do not|don'?t|dont|avoid|ensure|prefer|only|make sure|be sure|no)\b/i;
64
- function isEventRecord(rule) {
63
+ export const TITLE_OPENS_WITH_DIRECTIVE = /^\s*[-*+\d.\s]*(never|always|must|do not|don'?t|dont|avoid|ensure|prefer|only|make sure|be sure|no)\b/i;
64
+ export function isEventRecord(rule) {
65
65
  if (TITLE_OPENS_WITH_DIRECTIVE.test(rule.title))
66
66
  return false;
67
67
  if (IMPERATIVE_INSTRUCTION.test(rule.title))
@@ -126,7 +126,7 @@ function looksLikeBareCommand(text) {
126
126
  * Guarded by TITLE_OPENS_WITH_DIRECTIVE so a genuine prohibition whose body
127
127
  * is the forbidden command ("Never run: `rm -rf /`") is still a rule.
128
128
  */
129
- function isCommandDocumentation(rule) {
129
+ export function isCommandDocumentation(rule) {
130
130
  if (TITLE_OPENS_WITH_DIRECTIVE.test(rule.title))
131
131
  return false;
132
132
  if (FLAG_DOCUMENTATION.test(rule.text) || FLAG_DOCUMENTATION.test(rule.title))
@@ -1,5 +1,32 @@
1
1
  import type { TranscriptEvent, CheckResult } from "../types.js";
2
2
  import type { DeterministicClassification } from "./classify.js";
3
+ /**
4
+ * Appends the canonical long spelling of any short destructive flag the
5
+ * text uses, so a rule banning `git push --force` also catches `git push -f`.
6
+ *
7
+ * Extracted from searchHaystack 2026-09-15 so the pre-execution guard uses
8
+ * exactly the same aliasing as the post-hoc checker. A guard that missed
9
+ * `-f` while the report caught it would be worse than having neither.
10
+ */
11
+ export declare function canonicalise(text: string): string;
12
+ /**
13
+ * Word-boundary-aware match, not a naive substring search — otherwise a
14
+ * pattern like "git push --force" would false-positive on the SAFER
15
+ * "git push --force-with-lease" (caught by an actual failing test before
16
+ * this fix, not assumed).
17
+ *
18
+ * The trailing boundary only applies when the pattern itself ends in a
19
+ * word character — that's the only case where appending more word
20
+ * characters could form a genuinely different, longer token (like
21
+ * "--force" extending into "--force-with-lease"). A pattern that already
22
+ * ends in punctuation (e.g. "http://", ".env") can't be turned into a
23
+ * different token that way, and real occurrences of it (a real URL, a
24
+ * real filename) always have more characters immediately after — a real
25
+ * bug found by testing this against an actual "http://example.com"
26
+ * string: the old unconditional boundary made "http://" unmatchable
27
+ * against any real URL, ever.
28
+ */
29
+ export declare function matchesPattern(haystack: string, pattern: string): boolean;
3
30
  /**
4
31
  * Checks a single deterministic rule against the transcript by scanning
5
32
  * every event for the rule's literal banned pattern(s). No API calls, no
@@ -54,8 +54,19 @@ function searchHaystack(event) {
54
54
  const raw = eventSearchText(event);
55
55
  if (event.kind !== "tool_use")
56
56
  return raw;
57
- const extra = FLAG_ALIASES.filter((a) => a.context.test(raw) && a.short.test(raw)).map((a) => a.canonical);
58
- return extra.length > 0 ? `${raw} ${extra.join(" ")}` : raw;
57
+ return canonicalise(raw);
58
+ }
59
+ /**
60
+ * Appends the canonical long spelling of any short destructive flag the
61
+ * text uses, so a rule banning `git push --force` also catches `git push -f`.
62
+ *
63
+ * Extracted from searchHaystack 2026-09-15 so the pre-execution guard uses
64
+ * exactly the same aliasing as the post-hoc checker. A guard that missed
65
+ * `-f` while the report caught it would be worse than having neither.
66
+ */
67
+ export function canonicalise(text) {
68
+ const extra = FLAG_ALIASES.filter((a) => a.context.test(text) && a.short.test(text)).map((a) => a.canonical);
69
+ return extra.length > 0 ? `${text} ${extra.join(" ")}` : text;
59
70
  }
60
71
  function escapeRegex(literal) {
61
72
  return literal.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
@@ -77,7 +88,7 @@ function escapeRegex(literal) {
77
88
  * string: the old unconditional boundary made "http://" unmatchable
78
89
  * against any real URL, ever.
79
90
  */
80
- function matchesPattern(haystack, pattern) {
91
+ export function matchesPattern(haystack, pattern) {
81
92
  const lastChar = pattern[pattern.length - 1];
82
93
  const needsTrailingBoundary = /[\w-]/.test(lastChar);
83
94
  const suffix = needsTrailingBoundary ? "(?![\\w-])" : "";
@@ -0,0 +1,30 @@
1
+ /**
2
+ * Is this literal shaped like a command, rather than a noun?
3
+ *
4
+ * The single most important restriction in the guard, and it exists because
5
+ * the first version was measured before it shipped. Replaying 16,322 real
6
+ * tool calls against the 371 literal prohibitions in the corpus, it refused
7
+ * 62% of them. The literals doing the damage were ordinary words that rules
8
+ * name in passing — `browse`, `mix`, `index.ts`, `scripts:`, `AGENTS.md` —
9
+ * each of which appears in perfectly innocent commands all day.
10
+ *
11
+ * A prohibition worth blocking on names an invocation: `git push --force`,
12
+ * `rm -rf`, `npm publish`, or a bare flag like `--no-verify`. So: it must
13
+ * contain whitespace, or begin with a dash. A single bare word is never
14
+ * enough, which does mean a rule that says only "never use `rm`" is not
15
+ * enforced here. That is the right way round. The report still carries it,
16
+ * and refusing to run someone's command on the strength of one ambiguous
17
+ * word is not a trade worth making.
18
+ */
19
+ export declare function literalIsCommandShaped(literal: string): boolean;
20
+ /**
21
+ * Whether a command that is ABOUT TO RUN performs the forbidden thing.
22
+ *
23
+ * The post-hoc checker cannot answer this and says so: a literal in a
24
+ * transcript may be a violation, a grep, or an explanation, and nothing in
25
+ * the string distinguishes them. Before execution the question is narrower
26
+ * and mostly answerable, because the command is the act. What remains is
27
+ * separating the parts of a compound command that do something from the
28
+ * parts that only look at something, which is what this does.
29
+ */
30
+ export declare function commandRunsLiteral(command: string, literal: string): boolean;
@@ -0,0 +1,127 @@
1
+ import { withoutHeredocs } from "./shellCommand.js";
2
+ import { canonicalise, matchesPattern } from "./deterministicChecks.js";
3
+ /**
4
+ * Commands that read. A banned literal appearing as an argument to one of
5
+ * these is being searched for, printed, or paged — not run.
6
+ *
7
+ * This is the whole reason the post-hoc deterministic checker refuses to
8
+ * report a violation: a text match cannot tell an action from a mention.
9
+ * Before the tool runs, most of that ambiguity is gone — the command IS the
10
+ * action — but not all of it, because a command can still quote a literal
11
+ * while doing something harmless with it. This list is what is left of the
12
+ * problem, and it is deliberately a known list: anything not on it is
13
+ * treated as doing something.
14
+ *
15
+ * `sed` is absent on purpose. `sed -i` edits in place; plain `sed` does not,
16
+ * and the difference is handled below rather than by listing the name.
17
+ */
18
+ const READ_ONLY = new Set([
19
+ "grep", "rg", "ag", "ack", "egrep", "fgrep",
20
+ "echo", "printf", "cat", "bat", "head", "tail", "less", "more",
21
+ "find", "fd", "ls", "wc", "sort", "uniq", "diff", "comm",
22
+ "awk", "jq", "yq", "cut", "tr", "column", "tee",
23
+ "which", "type", "file", "stat", "man", "help",
24
+ ]);
25
+ /** Read-only git subcommands — `git log` cannot delete anything. */
26
+ const READ_ONLY_GIT = new Set(["log", "show", "diff", "status", "blame", "describe", "config", "remote", "branch", "tag", "ls-files", "rev-parse", "shortlog"]);
27
+ /**
28
+ * Splits a shell command into the pieces that run separately.
29
+ *
30
+ * Crude by design: this is not a shell parser and must never pretend to be
31
+ * one. It exists so that `grep "rm -rf" notes.txt && npm run build` is read
32
+ * as two things, one of which searches for a string and one of which does
33
+ * not, rather than as one blob containing a banned literal.
34
+ */
35
+ function segments(command) {
36
+ return withoutHeredocs(command)
37
+ .split(/\n|&&|\|\||[;|]/)
38
+ .map((s) => s.trim())
39
+ .filter((s) => s.length > 0);
40
+ }
41
+ /** The executable a segment invokes, with env assignments and `sudo` skipped. */
42
+ function leadingCommand(segment) {
43
+ const words = segment.split(/\s+/).filter(Boolean);
44
+ let i = 0;
45
+ while (i < words.length && (/^[A-Za-z_][A-Za-z0-9_]*=/.test(words[i]) || words[i] === "sudo" || words[i] === "command" || words[i] === "time"))
46
+ i++;
47
+ const exe = (words[i] ?? "").replace(/^.*\//, "");
48
+ return exe;
49
+ }
50
+ /**
51
+ * A commit message is text the user wrote, not a command being run.
52
+ *
53
+ * Real case in this project's own history: a commit message containing
54
+ * backticked shell examples. A message that quotes a banned command in order
55
+ * to describe it must not be treated as running it.
56
+ */
57
+ function withoutCommitMessage(segment) {
58
+ return segment.replace(/(-m|--message)(=|\s+)(['"])(?:\\.|(?!\3)[\s\S])*\3/g, "$1 <message>");
59
+ }
60
+ /**
61
+ * Does this segment actually DO the forbidden thing?
62
+ *
63
+ * Returns false for a segment that only reads, searches, or prints.
64
+ */
65
+ function segmentRunsLiteral(segment, literal) {
66
+ const exe = leadingCommand(segment);
67
+ if (READ_ONLY.has(exe))
68
+ return false;
69
+ if (exe === "sed" && !/\s-[A-Za-z]*i\b/.test(segment))
70
+ return false;
71
+ if (exe === "git") {
72
+ const sub = segment.split(/\s+/).filter(Boolean)[1] ?? "";
73
+ // A read-only git subcommand cannot be the destructive act — unless the
74
+ // rule names that subcommand itself, which is a real thing people write
75
+ // ("never run `git config user.email`"), so the literal still has to be
76
+ // checked against it. What is skipped is only the case where the literal
77
+ // is some OTHER command quoted inside a read-only one.
78
+ if (READ_ONLY_GIT.has(sub) && !literal.includes(`git ${sub}`))
79
+ return false;
80
+ }
81
+ return matchesPattern(canonicalise(withoutCommitMessage(segment)), literal);
82
+ }
83
+ /**
84
+ * Is this literal shaped like a command, rather than a noun?
85
+ *
86
+ * The single most important restriction in the guard, and it exists because
87
+ * the first version was measured before it shipped. Replaying 16,322 real
88
+ * tool calls against the 371 literal prohibitions in the corpus, it refused
89
+ * 62% of them. The literals doing the damage were ordinary words that rules
90
+ * name in passing — `browse`, `mix`, `index.ts`, `scripts:`, `AGENTS.md` —
91
+ * each of which appears in perfectly innocent commands all day.
92
+ *
93
+ * A prohibition worth blocking on names an invocation: `git push --force`,
94
+ * `rm -rf`, `npm publish`, or a bare flag like `--no-verify`. So: it must
95
+ * contain whitespace, or begin with a dash. A single bare word is never
96
+ * enough, which does mean a rule that says only "never use `rm`" is not
97
+ * enforced here. That is the right way round. The report still carries it,
98
+ * and refusing to run someone's command on the strength of one ambiguous
99
+ * word is not a trade worth making.
100
+ */
101
+ export function literalIsCommandShaped(literal) {
102
+ const t = literal.trim();
103
+ if (t.length === 0)
104
+ return false;
105
+ if (/^-{1,2}[A-Za-z]/.test(t))
106
+ return true;
107
+ // Must BEGIN like a command too, not merely contain a space: `, and` is a
108
+ // real corpus literal and it blocked 118 commands on its own.
109
+ if (!/^[A-Za-z][A-Za-z0-9_.\/-]*(\s|$)/.test(t))
110
+ return false;
111
+ return /\s/.test(t);
112
+ }
113
+ /**
114
+ * Whether a command that is ABOUT TO RUN performs the forbidden thing.
115
+ *
116
+ * The post-hoc checker cannot answer this and says so: a literal in a
117
+ * transcript may be a violation, a grep, or an explanation, and nothing in
118
+ * the string distinguishes them. Before execution the question is narrower
119
+ * and mostly answerable, because the command is the act. What remains is
120
+ * separating the parts of a compound command that do something from the
121
+ * parts that only look at something, which is what this does.
122
+ */
123
+ export function commandRunsLiteral(command, literal) {
124
+ if (!literalIsCommandShaped(literal))
125
+ return false;
126
+ return segments(command).some((s) => segmentRunsLiteral(s, literal));
127
+ }
package/dist/cli.js CHANGED
@@ -17,6 +17,7 @@ import { runCodeContentChecks } from "./checks/codeContent.js";
17
17
  import { runFileLifecycleChecks } from "./checks/fileLifecycle.js";
18
18
  import { runClaimEvidenceChecks } from "./checks/claimEvidence.js";
19
19
  import { runHook } from "./hook.js";
20
+ import { runGuard } from "./guard.js";
20
21
  import { runJudgmentChecks } from "./checks/judgmentChecks.js";
21
22
  import { generateReport, generateMarkdownReport } from "./report/generateReport.js";
22
23
  import { generateHtmlReport } from "./report/generateHtmlReport.js";
@@ -487,6 +488,53 @@ async function runRules(opts) {
487
488
  }
488
489
  return;
489
490
  }
491
+ /**
492
+ * Marking WHICH clause of a rule is the prohibition.
493
+ *
494
+ * The one thing a rules file never says. Blocking on every backtick in a
495
+ * forbidding rule refused 62.8% of 16,336 real tool calls, and the worst
496
+ * survivor after two narrowings refused `npm run build` 112 times against
497
+ * a rule that recommends it. So the guard blocks on nothing here until a
498
+ * person names the clause, and this is where they name it.
499
+ */
500
+ if (opts.forbid) {
501
+ const rule = findRule(opts.forbid);
502
+ if (!rule) {
503
+ console.error(`No rule in this project has the handle ${opts.forbid}.`);
504
+ console.error(`Handles come from \`rulereceipt check --show-skipped\`, and change if the rule's wording changes.`);
505
+ process.exitCode = 1;
506
+ return;
507
+ }
508
+ const literal = opts.literal?.trim();
509
+ if (!literal) {
510
+ console.error(`--forbid needs --literal "<the exact command this rule bans>".`);
511
+ console.error(`Copy it from the rule itself; it has to appear in the rule's text.`);
512
+ process.exitCode = 1;
513
+ return;
514
+ }
515
+ if (!`${rule.title}\n${rule.text ?? ""}`.includes(literal)) {
516
+ console.error(`That rule does not contain "${literal}".`);
517
+ console.error(`The mark has to name something the rule actually says, or a gate would`);
518
+ console.error(`refuse a command for a reason written nowhere.`);
519
+ process.exitCode = 1;
520
+ return;
521
+ }
522
+ const prior = overrides.get(opts.forbid)?.forbids ?? [];
523
+ const forbids = [...new Set([...prior, literal])];
524
+ saveOverride(cwd, {
525
+ hash: opts.forbid,
526
+ decision: overrides.get(opts.forbid)?.decision ?? "rule",
527
+ title: rule.title.replace(/\s+/g, " ").trim().slice(0, 200),
528
+ forbids,
529
+ });
530
+ console.log(`Saved to ${OVERRIDES_PATH}.`);
531
+ console.log(` "${rule.title.replace(/\s+/g, " ").trim().slice(0, 90)}"`);
532
+ console.log(` now blocks on: ${forbids.map((f) => `\`${f}\``).join(", ")}`);
533
+ console.log(`\nThis only takes effect if you run \`rulereceipt guard\` as a PreToolUse hook.`);
534
+ console.log(`Rewording the rule drops the mark, on purpose — it would otherwise carry`);
535
+ console.log(`your judgment onto words you never read.`);
536
+ return;
537
+ }
490
538
  if (opts.clear) {
491
539
  console.log(clearOverride(cwd, opts.clear) ? `Removed the correction for ${opts.clear}.` : `No correction stored for ${opts.clear}.`);
492
540
  return;
@@ -580,6 +628,12 @@ program
580
628
  .action(async () => {
581
629
  await runHook(needsLlmResult);
582
630
  });
631
+ program
632
+ .command("guard")
633
+ .description("run as a Claude Code PreToolUse hook - refuse a command that breaks a rule, before it runs (payload on stdin)")
634
+ .action(async () => {
635
+ await runGuard();
636
+ });
583
637
  program
584
638
  .command("doctor")
585
639
  .description("List every Claude Code hook and VS Code auto-task on this machine/project, flag anything suspicious")
@@ -597,6 +651,8 @@ program
597
651
  .description("correct what the classifier treats as a rule. Handles come from `check --show-skipped`.")
598
652
  .option("--include <handle>", "treat this item as a real rule and check it from now on")
599
653
  .option("--exclude <handle>", "treat this item as documentation and stop reporting it")
654
+ .option("--forbid <handle>", "mark which clause of this rule is the prohibition, so the guard may block on it")
655
+ .option("--literal <text>", "the exact banned command, used with --forbid; must appear in the rule")
600
656
  .option("--clear <handle>", "remove a stored correction")
601
657
  .option("--list", "show stored corrections (the default when no other flag is given)")
602
658
  .option("--coverage", "show which rules a configured hook might actually be enforcing, and which are prose only")
@@ -0,0 +1,42 @@
1
+ /**
2
+ * A PreToolUse hook that refuses a command before it runs.
3
+ *
4
+ * The Stop hook added on 2026-09-14 catches a finished session. That is too
5
+ * late for the case it most needs to cover: on 2026-04-25 a Cursor agent
6
+ * running Claude Opus 4.6 deleted PocketOS's production database and every
7
+ * volume-level backup in nine seconds, using a Railway token it found that
8
+ * had been created for managing domains. The agent had a rule — "NEVER run
9
+ * destructive/irreversible git commands...unless the user explicitly
10
+ * requests them" — and afterwards quoted it back, observing that what it had
11
+ * done was "far worse than a force push". A report would have described a
12
+ * database that was already gone.
13
+ *
14
+ * What it does NOT do is the thing it was built for. Blocking on a rule's
15
+ * banned command literal was measured before shipping and cut: see the note
16
+ * above `reason`. It refused 62% of 16,336 real commands, and the residue
17
+ * after two rounds of narrowing was still wrong in a way no matcher fixes.
18
+ * PocketOS would not have been stopped by this hook, and saying otherwise
19
+ * would be the exact failure this tool exists to catch.
20
+ *
21
+ * What remains is real and narrower: rules that name a FILE or a BRANCH.
22
+ * "Never modify `.env`", "never touch `migrations/`", "never commit to
23
+ * `main`". The classifier identified those as a path or a ref rather than
24
+ * guessing which backtick was the prohibition, so a refusal can be stated
25
+ * with a reason that holds up.
26
+ *
27
+ * Three properties, in the order they matter:
28
+ *
29
+ * 1. FORBIDDING rules only, answered by a structured checker. Never a
30
+ * judgment rule, never an LLM opinion, never a requirement, and never a
31
+ * bare command literal. Blocking someone's terminal on a guess is not a
32
+ * trade worth making at any hit rate.
33
+ *
34
+ * 2. It fails OPEN. Any error allows the command and writes to stderr. The
35
+ * opposite choice means a bug in this file stops someone from running
36
+ * anything at all, and they would remove the hook within the hour — which
37
+ * leaves them with no guard rather than an imperfect one.
38
+ *
39
+ * 3. It says which rule and why, in the refusal itself, because a block with
40
+ * no reason is indistinguishable from a broken tool.
41
+ */
42
+ export declare function runGuard(): Promise<void>;
package/dist/guard.js ADDED
@@ -0,0 +1,219 @@
1
+ import { loadRules } from "./rules.js";
2
+ import { classifyRules } from "./checks/classify.js";
3
+ import { runCodeContentChecks } from "./checks/codeContent.js";
4
+ import { runFileLifecycleChecks } from "./checks/fileLifecycle.js";
5
+ import { runGitBranchPolicyChecks } from "./checks/gitBranchPolicy.js";
6
+ import { loadOverrides, ruleFingerprint, ratifiedForbids } from "./overrides.js";
7
+ import { commandRunsLiteral } from "./checks/proposedAction.js";
8
+ function readStdin() {
9
+ return new Promise((resolve) => {
10
+ let data = "";
11
+ if (process.stdin.isTTY)
12
+ return resolve("");
13
+ process.stdin.setEncoding("utf8");
14
+ process.stdin.on("data", (c) => (data += c));
15
+ process.stdin.on("end", () => resolve(data));
16
+ process.stdin.on("error", () => resolve(""));
17
+ });
18
+ }
19
+ /**
20
+ * The rules this can answer BEFORE the command runs.
21
+ *
22
+ * Deliberately only the forbidding ones. "Always run the tests before
23
+ * committing" cannot be judged from a single proposed call — whether it was
24
+ * already satisfied is a fact about the session, not about this command —
25
+ * and a guard that blocked on it would fire on the first commit of every
26
+ * session. Requirements stay with the report and the Stop hook.
27
+ */
28
+ function forbidRules(cwd) {
29
+ const overrides = loadOverrides(cwd);
30
+ return classifyRules(loadRules(cwd)).filter((c) => {
31
+ if (overrides.get(ruleFingerprint(c.rule))?.decision === "notARule")
32
+ return false;
33
+ return "polarity" in c && c.polarity === "forbid";
34
+ });
35
+ }
36
+ /**
37
+ * Runs the structured checkers against a single PROPOSED action.
38
+ *
39
+ * They already take a list of events and ask what it did, so a one-event
40
+ * list describing what is about to happen is exactly the right input. This
41
+ * is why the guard cannot drift from the report: same checkers, same
42
+ * verdicts, different tense.
43
+ */
44
+ function structuredBlocks(cwd, event) {
45
+ const cls = forbidRules(cwd);
46
+ const of = (k) => cls.filter((c) => c.kind === k);
47
+ const results = [
48
+ ...runCodeContentChecks(of("codeContent"), [event]),
49
+ ...runFileLifecycleChecks(of("fileLifecycle"), [event]),
50
+ ...runGitBranchPolicyChecks(of("gitBranchPolicy"), [event]),
51
+ ];
52
+ return results
53
+ .filter((r) => r.status === "FAIL")
54
+ .map((r) => ({
55
+ rule: { id: r.ruleId, title: r.ruleTitle, text: "", source: r.ruleSource },
56
+ why: r.evidence,
57
+ }));
58
+ }
59
+ /**
60
+ * A command ban blocks only where a person marked the clause.
61
+ *
62
+ * The unratified version was measured before shipping and cut: blocking on a
63
+ * rule's command literals refused 62.8% of 16,336 real tool calls. Two
64
+ * narrowings reached 2.49% and the residue had no matcher fix — a rule
65
+ * titled "Feature Validation" refused `npm run build` 112 times, because it
66
+ * forbids running Playwright unprompted and RECOMMENDS the build command,
67
+ * which is its only command-shaped literal. Another refused plain
68
+ * `git status`, its backticks holding both the ban and the alternative.
69
+ *
70
+ * Nothing in a rules file marks which backtick is the prohibition. So this
71
+ * path reads only what someone declared, and yields nothing otherwise. The
72
+ * declaration is keyed on the rule's content hash and re-checked against the
73
+ * rule's current text, so a reworded rule loses its mark rather than
74
+ * carrying a judgement onto words nobody read.
75
+ *
76
+ * The important property is what happens by default: an unmarked rule cannot
77
+ * block, at any confidence, ever. That is the whole difference between this
78
+ * and the version that refused two thirds of everything.
79
+ */
80
+ function ratifiedLiteralBlocks(cwd, command) {
81
+ const overrides = loadOverrides(cwd);
82
+ const blocks = [];
83
+ for (const c of forbidRules(cwd)) {
84
+ if (c.kind !== "deterministic")
85
+ continue;
86
+ for (const literal of ratifiedForbids(overrides, c.rule)) {
87
+ if (!commandRunsLiteral(command, literal))
88
+ continue;
89
+ blocks.push({ rule: c.rule, why: `the command about to run does \`${literal}\`, which this rule forbids (marked by you, not inferred)` });
90
+ break;
91
+ }
92
+ }
93
+ return blocks;
94
+ }
95
+ /**
96
+ * Literal command bans do NOT block unless ratified. Measured, then cut.
97
+ *
98
+ * This was the point of the feature and it does not survive its own
99
+ * measurement. Replaying 16,336 real tool calls against every forbidding
100
+ * rule in the corpus, blocking on literals refused 62% of commands. Two
101
+ * restrictions brought that to 1.6% — the literal has to be shaped like an
102
+ * invocation rather than a noun, and the rule has to name exactly one of
103
+ * them — and what remained was still wrong in a way no matcher can fix.
104
+ *
105
+ * A rule titled "Feature Validation" refused `npm run build` 112 times. It
106
+ * forbids running Playwright without asking; it RECOMMENDS `npm run build`.
107
+ * Forbid polarity, one command-shaped literal, and that literal is the
108
+ * approved command. Another rule refused plain `git status`, because its
109
+ * backticks hold both the thing it bans and the thing it suggests instead.
110
+ *
111
+ * Nothing in a rules file marks which backtick is the prohibition. The
112
+ * report can live with that — it says UNCLEAR and a person reads it. A
113
+ * blocker cannot: it would refuse the recommended command with a confident
114
+ * explanation. So command bans stay in the report and the Stop hook, and
115
+ * only the structured checkers, which know what kind of thing they are
116
+ * looking at because the classifier identified a path or a branch, can
117
+ * refuse anything here.
118
+ *
119
+ * The way back to command bans is an explicit opt-in — the user naming which
120
+ * rules may block — not a cleverer guess. That is a design question for
121
+ * whoever asks for it, not a default.
122
+ */
123
+ function reason(blocks) {
124
+ const lines = blocks.map((b) => ` • Rule ${b.rule.id} — ${b.rule.title}\n ${b.why}`);
125
+ const n = blocks.length;
126
+ return (`RuleReceipt blocked this: it breaks ${n === 1 ? "a rule" : `${n} rules`} in CLAUDE.md.\n\n` +
127
+ lines.join("\n\n") +
128
+ `\n\nIf the rule should not apply here, say so to the user and let them decide. Do not work around the rule by rephrasing the command.`);
129
+ }
130
+ /**
131
+ * A PreToolUse hook that refuses a command before it runs.
132
+ *
133
+ * The Stop hook added on 2026-09-14 catches a finished session. That is too
134
+ * late for the case it most needs to cover: on 2026-04-25 a Cursor agent
135
+ * running Claude Opus 4.6 deleted PocketOS's production database and every
136
+ * volume-level backup in nine seconds, using a Railway token it found that
137
+ * had been created for managing domains. The agent had a rule — "NEVER run
138
+ * destructive/irreversible git commands...unless the user explicitly
139
+ * requests them" — and afterwards quoted it back, observing that what it had
140
+ * done was "far worse than a force push". A report would have described a
141
+ * database that was already gone.
142
+ *
143
+ * What it does NOT do is the thing it was built for. Blocking on a rule's
144
+ * banned command literal was measured before shipping and cut: see the note
145
+ * above `reason`. It refused 62% of 16,336 real commands, and the residue
146
+ * after two rounds of narrowing was still wrong in a way no matcher fixes.
147
+ * PocketOS would not have been stopped by this hook, and saying otherwise
148
+ * would be the exact failure this tool exists to catch.
149
+ *
150
+ * What remains is real and narrower: rules that name a FILE or a BRANCH.
151
+ * "Never modify `.env`", "never touch `migrations/`", "never commit to
152
+ * `main`". The classifier identified those as a path or a ref rather than
153
+ * guessing which backtick was the prohibition, so a refusal can be stated
154
+ * with a reason that holds up.
155
+ *
156
+ * Three properties, in the order they matter:
157
+ *
158
+ * 1. FORBIDDING rules only, answered by a structured checker. Never a
159
+ * judgment rule, never an LLM opinion, never a requirement, and never a
160
+ * bare command literal. Blocking someone's terminal on a guess is not a
161
+ * trade worth making at any hit rate.
162
+ *
163
+ * 2. It fails OPEN. Any error allows the command and writes to stderr. The
164
+ * opposite choice means a bug in this file stops someone from running
165
+ * anything at all, and they would remove the hook within the hour — which
166
+ * leaves them with no guard rather than an imperfect one.
167
+ *
168
+ * 3. It says which rule and why, in the refusal itself, because a block with
169
+ * no reason is indistinguishable from a broken tool.
170
+ */
171
+ export async function runGuard() {
172
+ const allow = () => {
173
+ process.stdout.write(JSON.stringify({}));
174
+ };
175
+ try {
176
+ const raw = await readStdin();
177
+ const input = raw ? JSON.parse(raw) : {};
178
+ const cwd = input.cwd || process.cwd();
179
+ const tool = input.tool_name ?? "";
180
+ const toolInput = input.tool_input ?? {};
181
+ if (loadRules(cwd).length === 0)
182
+ return allow();
183
+ let blocks = [];
184
+ if (tool === "Bash" && typeof toolInput.command === "string") {
185
+ const event = {
186
+ role: "assistant", kind: "tool_use", toolName: "Bash",
187
+ input: { command: toolInput.command }, timestamp: "",
188
+ };
189
+ blocks = [...structuredBlocks(cwd, event), ...ratifiedLiteralBlocks(cwd, toolInput.command)];
190
+ }
191
+ else if (tool === "Write" || tool === "Edit" || tool === "NotebookEdit") {
192
+ const event = {
193
+ role: "assistant", kind: "tool_use", toolName: tool,
194
+ input: toolInput, timestamp: "",
195
+ };
196
+ blocks = structuredBlocks(cwd, event);
197
+ }
198
+ else {
199
+ return allow();
200
+ }
201
+ if (blocks.length === 0)
202
+ return allow();
203
+ process.stdout.write(JSON.stringify({
204
+ hookSpecificOutput: {
205
+ hookEventName: "PreToolUse",
206
+ permissionDecision: "deny",
207
+ permissionDecisionReason: reason(blocks),
208
+ },
209
+ }));
210
+ // Exit 2 is what actually blocks the call; the JSON carries the reason.
211
+ process.exitCode = 2;
212
+ }
213
+ catch (err) {
214
+ // Fail open, and never with exit 2 — an exit 2 from a crash would block
215
+ // every command the session tries to run.
216
+ process.stderr.write(`rulereceipt guard: allowing command, check did not complete (${err instanceof Error ? err.message : String(err)})\n`);
217
+ allow();
218
+ }
219
+ }
package/dist/hook.js CHANGED
@@ -22,12 +22,25 @@ function readStdin() {
22
22
  * what to do and repeating it invites the model to argue with the wording
23
23
  * instead of going and running the thing.
24
24
  */
25
- function blockReason(failures) {
26
- const lines = failures.map((f) => ` • Rule ${f.ruleId} — ${f.ruleTitle}\n ${f.evidence}`);
27
- const n = failures.length;
28
- return (`RuleReceipt: ${n} rule${n === 1 ? "" : "s"} in CLAUDE.md ${n === 1 ? "was" : "were"} not followed in this session.\n\n` +
29
- lines.join("\n\n") +
30
- `\n\nDo not report this work as finished until the above is resolved or you have said plainly, to the user, that it is still open and why.`);
25
+ function blockReason(failures, unverified) {
26
+ const parts = [];
27
+ if (failures.length > 0) {
28
+ const n = failures.length;
29
+ parts.push(`RuleReceipt: ${n} rule${n === 1 ? "" : "s"} in CLAUDE.md ${n === 1 ? "was" : "were"} not followed in this session.\n\n` +
30
+ failures.map((f) => ` • Rule ${f.ruleId} — ${f.ruleTitle}\n ${f.evidence}`).join("\n\n"));
31
+ }
32
+ // Worded as a question about evidence rather than as a finding, because
33
+ // that is what it is. Nothing here says the claim is false. It says
34
+ // nothing in this session shows it to be true, and that the difference
35
+ // belongs to the user rather than to the summary.
36
+ if (unverified.length > 0) {
37
+ parts.push(`RuleReceipt: this session claims work is done, and nothing recorded here verifies it.\n\n` +
38
+ unverified.map((u) => ` • Rule ${u.ruleId} — ${u.ruleTitle}\n ${u.evidence}`).join("\n\n") +
39
+ `\n\nThis is not a claim that you are wrong. It may well have been verified somewhere this transcript cannot see.`);
40
+ }
41
+ parts.push(`Do not report this work as finished until the above is resolved, or you have said plainly, to the user, ` +
42
+ `what was actually verified and what was not.`);
43
+ return parts.join("\n\n");
31
44
  }
32
45
  /**
33
46
  * A Stop hook that refuses to let a session end on a broken rule.
@@ -88,9 +101,13 @@ export async function runHook(needsLlmResult) {
88
101
  // it is part of the contract.
89
102
  const { results } = await evaluateSession(cwd, rules, events, false, needsLlmResult);
90
103
  const failures = results.filter((r) => r.status === "FAIL" && r.outcome !== "not_run");
91
- if (failures.length === 0)
104
+ // Deliberately NOT a FAIL. See CheckResult.unverifiedClaim: the report
105
+ // calls this unclear and the gate refuses it, and that is the only place
106
+ // the two are allowed to disagree.
107
+ const unverified = results.filter((r) => r.unverifiedClaim === true);
108
+ if (failures.length === 0 && unverified.length === 0)
92
109
  return void emit({});
93
- return void emit({ decision: "block", reason: blockReason(failures) });
110
+ return void emit({ decision: "block", reason: blockReason(failures, unverified) });
94
111
  }
95
112
  catch (err) {
96
113
  // Property 3: fail open, but never silently.
@@ -26,6 +26,21 @@ export interface Override {
26
26
  decision: Decision;
27
27
  /** Stored for humans reading the file, and to explain a stale entry. */
28
28
  title: string;
29
+ /**
30
+ * The literal(s) a person has marked as the PROHIBITION in this rule.
31
+ *
32
+ * A rules file does not say which of its backticks is the thing being
33
+ * banned. Measured: blocking on all of them refused 62.8% of 16,336 real
34
+ * tool calls, and the residue after two narrowings still refused
35
+ * `npm run build` 112 times, because the rule that named it forbids
36
+ * running Playwright and RECOMMENDS the build command. No matcher fixes
37
+ * that — the information is not in the text.
38
+ *
39
+ * So it is declared, once, and only what is declared can block. Absent
40
+ * means absent: there is deliberately no fallback to "probably the first
41
+ * literal", because that fallback is the bug.
42
+ */
43
+ forbids?: string[];
29
44
  }
30
45
  export declare const OVERRIDES_PATH: string;
31
46
  /**
@@ -60,3 +75,21 @@ export declare function clearOverride(cwd: string, hash: string): boolean;
60
75
  * correction is no longer being applied instead of assuming it still is.
61
76
  */
62
77
  export declare function staleOverrides(overrides: Map<string, Override>, rules: Rule[]): Override[];
78
+ /**
79
+ * The prohibition literals a person has declared for this rule, or none.
80
+ *
81
+ * Two conditions, both required, and both are refusals to infer:
82
+ *
83
+ * 1. The mark is keyed on the rule's CONTENT hash, so rewording the rule
84
+ * drops it rather than reattaching a judgement to text nobody read. Same
85
+ * reasoning as ruleFingerprint, one field down.
86
+ *
87
+ * 2. The literal must still appear in the rule. A mark that survives a
88
+ * partial edit and names something the rule no longer mentions would
89
+ * block on a phrase with nothing behind it — a gate refusing a command
90
+ * for a reason that is no longer written anywhere, which is the worst
91
+ * failure available to a gate.
92
+ *
93
+ * Returns [] for anything unratified. There is no fallback on purpose.
94
+ */
95
+ export declare function ratifiedForbids(overrides: Map<string, Override>, rule: Rule): string[];
package/dist/overrides.js CHANGED
@@ -32,7 +32,15 @@ export function loadOverrides(cwd) {
32
32
  return map;
33
33
  for (const o of parsed.overrides) {
34
34
  if (typeof o?.hash === "string" && (o.decision === "rule" || o.decision === "notARule")) {
35
- map.set(o.hash, { hash: o.hash, decision: o.decision, title: String(o.title ?? "") });
35
+ const forbids = Array.isArray(o.forbids)
36
+ ? o.forbids.filter((x) => typeof x === "string" && x.trim().length > 0)
37
+ : undefined;
38
+ map.set(o.hash, {
39
+ hash: o.hash,
40
+ decision: o.decision,
41
+ title: String(o.title ?? ""),
42
+ ...(forbids && forbids.length > 0 ? { forbids } : {}),
43
+ });
36
44
  }
37
45
  }
38
46
  }
@@ -83,3 +91,27 @@ export function staleOverrides(overrides, rules) {
83
91
  const live = new Set(rules.map(ruleFingerprint));
84
92
  return [...overrides.values()].filter((o) => !live.has(o.hash));
85
93
  }
94
+ /**
95
+ * The prohibition literals a person has declared for this rule, or none.
96
+ *
97
+ * Two conditions, both required, and both are refusals to infer:
98
+ *
99
+ * 1. The mark is keyed on the rule's CONTENT hash, so rewording the rule
100
+ * drops it rather than reattaching a judgement to text nobody read. Same
101
+ * reasoning as ruleFingerprint, one field down.
102
+ *
103
+ * 2. The literal must still appear in the rule. A mark that survives a
104
+ * partial edit and names something the rule no longer mentions would
105
+ * block on a phrase with nothing behind it — a gate refusing a command
106
+ * for a reason that is no longer written anywhere, which is the worst
107
+ * failure available to a gate.
108
+ *
109
+ * Returns [] for anything unratified. There is no fallback on purpose.
110
+ */
111
+ export function ratifiedForbids(overrides, rule) {
112
+ const entry = overrides.get(ruleFingerprint(rule));
113
+ if (!entry?.forbids)
114
+ return [];
115
+ const body = `${rule.title}\n${rule.text ?? ""}`;
116
+ return entry.forbids.filter((literal) => body.includes(literal));
117
+ }
@@ -117,7 +117,7 @@ export function parseLine(line) {
117
117
  events.push({ role: "assistant", kind: "text", text: b.text, timestamp });
118
118
  }
119
119
  else if (b.type === "tool_use" && typeof b.name === "string") {
120
- events.push({ role: "assistant", kind: "tool_use", toolName: b.name, input: b.input, timestamp });
120
+ events.push({ role: "assistant", kind: "tool_use", toolName: b.name, input: b.input, timestamp, toolUseId: typeof b.id === "string" ? b.id : undefined });
121
121
  }
122
122
  }
123
123
  }
@@ -140,6 +140,7 @@ export function parseLine(line) {
140
140
  content: extractToolResultText(b.content),
141
141
  isError: b.is_error === true,
142
142
  timestamp,
143
+ toolUseId: typeof b.tool_use_id === "string" ? b.tool_use_id : undefined,
143
144
  });
144
145
  }
145
146
  }
package/dist/types.d.ts CHANGED
@@ -16,6 +16,17 @@ export interface TranscriptToolUseEvent {
16
16
  toolName: string;
17
17
  input: unknown;
18
18
  timestamp: string;
19
+ /**
20
+ * The tool_use id, when the transcript carries one.
21
+ *
22
+ * Added 2026-09-16. Without it a result can only be attributed to "the
23
+ * call immediately before", which was documented here as safe on the
24
+ * grounds that every turn contained exactly one tool call. Parallel tool
25
+ * calls are now ordinary — a real session issues three in a row and then
26
+ * one result — and under that shape the positional rule silently drops
27
+ * the test run that a completion claim depended on.
28
+ */
29
+ toolUseId?: string;
19
30
  }
20
31
  export interface TranscriptToolResultEvent {
21
32
  role: "user";
@@ -23,6 +34,8 @@ export interface TranscriptToolResultEvent {
23
34
  content: string;
24
35
  isError: boolean;
25
36
  timestamp: string;
37
+ /** The id of the call this result belongs to, when the transcript has it. */
38
+ toolUseId?: string;
26
39
  }
27
40
  export type TranscriptEvent = TranscriptTextEvent | TranscriptToolUseEvent | TranscriptToolResultEvent;
28
41
  export type CheckStatus = "PASS" | "FAIL" | "UNCLEAR";
@@ -74,6 +87,23 @@ export interface CheckResult {
74
87
  * not blur them.
75
88
  */
76
89
  needsHuman?: boolean;
90
+ /**
91
+ * A completion claim that nothing in this session verified.
92
+ *
93
+ * The one place the report and the gate deliberately disagree. The report
94
+ * renders this UNCLEAR, because absence genuinely proves nothing — the
95
+ * tests may have run in another terminal and a transcript cannot see that.
96
+ * The Stop hook refuses the exit on it anyway, because it is not asserting
97
+ * a violation: it is declining to let "done" end a session when nothing
98
+ * here backs it, and the way out is one sentence saying so out loud.
99
+ *
100
+ * Added 2026-09-16. It is the case anthropics/claude-code#90542 is about
101
+ * from end to end — deploy artifacts produced and the application never
102
+ * opened, a completion whose evidence is a plan, "state what was VERIFIED
103
+ * versus only PLANNED" — and the gate was silent on all of it, firing only
104
+ * when a recorded run CONTRADICTED a claim.
105
+ */
106
+ unverifiedClaim?: boolean;
77
107
  /**
78
108
  * The outcome in the five-value vocabulary. Optional while the checkers
79
109
  * are migrated one at a time; `status` remains the fallback.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "rulereceipt",
3
- "version": "0.1.37",
3
+ "version": "0.1.39",
4
4
  "description": "Checks whether a Claude Code session actually followed your CLAUDE.md / AGENTS.md rules, with evidence.",
5
5
  "repository": {
6
6
  "type": "git",