@codyswann/lisa 2.314.0 → 2.316.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (86) hide show
  1. package/dist/core/upstream-evidence-manifest.d.ts.map +1 -1
  2. package/dist/core/upstream-evidence-manifest.js +26 -4
  3. package/dist/core/upstream-evidence-manifest.js.map +1 -1
  4. package/package.json +1 -1
  5. package/plugins/lisa/.claude-plugin/plugin.json +1 -1
  6. package/plugins/lisa/.codex-plugin/plugin.json +1 -1
  7. package/plugins/lisa/hooks/enforce-verification-gate.sh +52 -2
  8. package/plugins/lisa/rules/eager/automation-runbook-contract.md +25 -0
  9. package/plugins/lisa/rules/eager/empirical-inquiry.md +1 -0
  10. package/plugins/lisa/rules/eager/settled-decisions.md +55 -0
  11. package/plugins/lisa/rules/eager/stale-state-claims.md +26 -0
  12. package/plugins/lisa/rules/reference/falsifiable-checks.md +1 -0
  13. package/plugins/lisa/rules/reference/settled-decisions.md +29 -0
  14. package/plugins/lisa/rules/reference/stale-state-claims.md +88 -0
  15. package/plugins/lisa-agy/plugin.json +1 -1
  16. package/plugins/lisa-cdk/.claude-plugin/plugin.json +1 -1
  17. package/plugins/lisa-cdk/.codex-plugin/plugin.json +1 -1
  18. package/plugins/lisa-cdk-agy/plugin.json +1 -1
  19. package/plugins/lisa-cdk-copilot/.claude-plugin/plugin.json +1 -1
  20. package/plugins/lisa-cdk-cursor/.claude-plugin/plugin.json +1 -1
  21. package/plugins/lisa-copilot/.claude-plugin/plugin.json +1 -1
  22. package/plugins/lisa-copilot/hooks/enforce-verification-gate.sh +52 -2
  23. package/plugins/lisa-copilot/rules/eager/automation-runbook-contract.md +25 -0
  24. package/plugins/lisa-copilot/rules/eager/empirical-inquiry.md +1 -0
  25. package/plugins/lisa-copilot/rules/eager/settled-decisions.md +55 -0
  26. package/plugins/lisa-copilot/rules/eager/stale-state-claims.md +26 -0
  27. package/plugins/lisa-copilot/rules/reference/falsifiable-checks.md +1 -0
  28. package/plugins/lisa-copilot/rules/reference/settled-decisions.md +29 -0
  29. package/plugins/lisa-copilot/rules/reference/stale-state-claims.md +88 -0
  30. package/plugins/lisa-cursor/.claude-plugin/plugin.json +1 -1
  31. package/plugins/lisa-cursor/hooks/enforce-verification-gate.sh +52 -2
  32. package/plugins/lisa-cursor/rules/automation-runbook-contract.mdc +25 -0
  33. package/plugins/lisa-cursor/rules/empirical-inquiry.mdc +1 -0
  34. package/plugins/lisa-cursor/rules/falsifiable-checks-reference.mdc +1 -0
  35. package/plugins/lisa-cursor/rules/settled-decisions-reference.mdc +34 -0
  36. package/plugins/lisa-cursor/rules/settled-decisions.mdc +60 -0
  37. package/plugins/lisa-cursor/rules/stale-state-claims-reference.mdc +93 -0
  38. package/plugins/lisa-cursor/rules/stale-state-claims.mdc +31 -0
  39. package/plugins/lisa-expo/.claude-plugin/plugin.json +1 -1
  40. package/plugins/lisa-expo/.codex-plugin/plugin.json +1 -1
  41. package/plugins/lisa-expo-agy/plugin.json +1 -1
  42. package/plugins/lisa-expo-copilot/.claude-plugin/plugin.json +1 -1
  43. package/plugins/lisa-expo-cursor/.claude-plugin/plugin.json +1 -1
  44. package/plugins/lisa-harper-fabric/.claude-plugin/plugin.json +1 -1
  45. package/plugins/lisa-harper-fabric/.codex-plugin/plugin.json +1 -1
  46. package/plugins/lisa-harper-fabric-agy/plugin.json +1 -1
  47. package/plugins/lisa-harper-fabric-copilot/.claude-plugin/plugin.json +1 -1
  48. package/plugins/lisa-harper-fabric-cursor/.claude-plugin/plugin.json +1 -1
  49. package/plugins/lisa-nestjs/.claude-plugin/plugin.json +1 -1
  50. package/plugins/lisa-nestjs/.codex-plugin/plugin.json +1 -1
  51. package/plugins/lisa-nestjs-agy/plugin.json +1 -1
  52. package/plugins/lisa-nestjs-copilot/.claude-plugin/plugin.json +1 -1
  53. package/plugins/lisa-nestjs-cursor/.claude-plugin/plugin.json +1 -1
  54. package/plugins/lisa-openclaw/.claude-plugin/plugin.json +1 -1
  55. package/plugins/lisa-openclaw/.codex-plugin/plugin.json +1 -1
  56. package/plugins/lisa-openclaw-agy/plugin.json +1 -1
  57. package/plugins/lisa-openclaw-copilot/.claude-plugin/plugin.json +1 -1
  58. package/plugins/lisa-openclaw-cursor/.claude-plugin/plugin.json +1 -1
  59. package/plugins/lisa-phaser/.claude-plugin/plugin.json +1 -1
  60. package/plugins/lisa-phaser/.codex-plugin/plugin.json +1 -1
  61. package/plugins/lisa-phaser-agy/plugin.json +1 -1
  62. package/plugins/lisa-phaser-copilot/.claude-plugin/plugin.json +1 -1
  63. package/plugins/lisa-phaser-cursor/.claude-plugin/plugin.json +1 -1
  64. package/plugins/lisa-rails/.claude-plugin/plugin.json +1 -1
  65. package/plugins/lisa-rails/.codex-plugin/plugin.json +1 -1
  66. package/plugins/lisa-rails-agy/plugin.json +1 -1
  67. package/plugins/lisa-rails-copilot/.claude-plugin/plugin.json +1 -1
  68. package/plugins/lisa-rails-cursor/.claude-plugin/plugin.json +1 -1
  69. package/plugins/lisa-typescript/.claude-plugin/plugin.json +1 -1
  70. package/plugins/lisa-typescript/.codex-plugin/plugin.json +1 -1
  71. package/plugins/lisa-typescript-agy/plugin.json +1 -1
  72. package/plugins/lisa-typescript-copilot/.claude-plugin/plugin.json +1 -1
  73. package/plugins/lisa-typescript-cursor/.claude-plugin/plugin.json +1 -1
  74. package/plugins/lisa-wiki/.claude-plugin/plugin.json +1 -1
  75. package/plugins/lisa-wiki/.codex-plugin/plugin.json +1 -1
  76. package/plugins/lisa-wiki-agy/plugin.json +1 -1
  77. package/plugins/lisa-wiki-copilot/.claude-plugin/plugin.json +1 -1
  78. package/plugins/lisa-wiki-cursor/.claude-plugin/plugin.json +1 -1
  79. package/plugins/src/base/hooks/enforce-verification-gate.sh +52 -2
  80. package/plugins/src/base/rules/eager/automation-runbook-contract.md +25 -0
  81. package/plugins/src/base/rules/eager/empirical-inquiry.md +1 -0
  82. package/plugins/src/base/rules/eager/settled-decisions.md +55 -0
  83. package/plugins/src/base/rules/eager/stale-state-claims.md +26 -0
  84. package/plugins/src/base/rules/reference/falsifiable-checks.md +1 -0
  85. package/plugins/src/base/rules/reference/settled-decisions.md +29 -0
  86. package/plugins/src/base/rules/reference/stale-state-claims.md +88 -0
@@ -0,0 +1,55 @@
1
+ # Settled Decisions (load-bearing)
2
+
3
+ **Never ask the operator a question that a standing preference or your own gathered evidence has
4
+ already answered.** Re-asking a settled decision is not caution — it hands back a choice the
5
+ operator already made and makes them make it twice.
6
+
7
+ ## The test
8
+
9
+ Before asking anything, check the three sources that may already hold the answer:
10
+
11
+ 1. **A standing preference** — something the operator has told you once and expects to hold:
12
+ recorded in project memory, `.lisa.config.json`, a project rule, `CLAUDE.md`/`AGENTS.md`, or
13
+ stated earlier in this conversation. "Every PR gets auto-merge and gets watched to merge" is a
14
+ standing preference; asking "want me to watch this PR?" violates it.
15
+ 2. **Evidence you already gathered** — if the research, measurement or code read you just performed
16
+ resolves the question, the question is answered. Report the decision and the evidence for it.
17
+ 3. **A convention with an obvious default** — where one option is clearly conventional and the other
18
+ needs a reason, take the conventional one and say so in one line.
19
+
20
+ If any source answers it: **act, and state the decision in a clause.** Do not convert it into a
21
+ question.
22
+
23
+ ## When asking IS right
24
+
25
+ Ask when the answer would change the work *and* you genuinely cannot derive it:
26
+
27
+ - The options lead to materially different deliverables and nothing in scope decides between them.
28
+ - Proceeding on a wrong assumption would be unsafe, destructive, or waste substantial work.
29
+ - The answer is a human judgement the artifacts do not contain — a product decision, a design
30
+ vocabulary that does not exist yet, a risk the operator owns.
31
+
32
+ That kind of question is load-bearing and should be asked plainly, once, with a recommendation.
33
+
34
+ ## The failure mode this rule exists to stop
35
+
36
+ Asking feels collaborative, so it gets over-applied — especially at the end of a report, where a
37
+ trailing question reads as deference. It is not deference when the answer was already given; it is
38
+ work handed back. Two specific tells:
39
+
40
+ - **Half-applying an instruction.** Doing the first half of a standing preference automatically and
41
+ asking permission for the second half of the same preference.
42
+ - **Asking after the research answered it.** Completing an investigation that points one direction
43
+ unambiguously, then presenting the conclusion as an open question.
44
+
45
+ ## Partial application is worse than either extreme
46
+
47
+ If a standing preference covers a multi-step behavior, apply **all** of it. Applying part and asking
48
+ about the rest produces the worst outcome: the operator is interrupted *and* the instruction was not
49
+ honored.
50
+
51
+ ## Recording, not asking
52
+
53
+ When you take a settled decision, make it auditable in one clause — "auto-merge on, per your standing
54
+ preference", "scoped app-wide, because the Android finding decides it". That gives the operator the
55
+ chance to correct it without requiring them to answer first.
@@ -0,0 +1,26 @@
1
+ # Stale State Claims — "Not Yet" Expires, the Note Does Not (load-bearing)
2
+
3
+ A comment, docstring, config annotation, or work-item note that records a **temporary** state — "not yet", "pending", "waiting on X", "human-gated", "will flip once Y ships" — was accurate the day it was written and is believed long after it stopped being true. Nobody revisits prose when the condition clears.
4
+
5
+ It then misdirects with full authority: work gets skipped as blocked when nothing blocks it, re-planned as undone when it already shipped, or escalated to a person who made that exact decision weeks ago.
6
+
7
+ This is the sibling of `falsifiable-checks`. That rule is about **instruments that cannot fail**; this one is about **assertions of state that have expired**. Both read as authoritative and neither announces its own decay.
8
+
9
+ ## The four ways a recorded state outlives its truth
10
+
11
+ Each has been observed; each was believed long after it went false:
12
+
13
+ 1. **Expired blocker note** — prose asserting a not-yet condition that has since cleared. An endpoint left unset under a comment saying its routing was "not yet promoted" had been serving live for days; an agent planned work to enable what was already enabled.
14
+ 2. **Stale gate marker** — a "pending human approval / not provisioned yet" annotation whose precondition is satisfied. One held for twelve days after the thing it waited on existed, and the decision was escalated to a person surprised to be asked.
15
+ 3. **Prediction buried by closure** — a warning recorded in a comment on an item that then closes. Closure deletes the warning: one correctly predicted a defect, the item closed 25 minutes later, and the defect sat unnoticed for eight days while a dependent item was parked waiting on it.
16
+ 4. **Silent waiting gate** — a queue or approval state with no surfacing mechanism, which holds work for exactly as long as nobody happens to look. One held a security-relevant fix undeployed for two days.
17
+
18
+ ## Mandatory
19
+
20
+ - **A recorded blocker is a claim about the past — check the present before acting on it.** Before planning around it, escalating it, or reporting it as a blocker, probe the live state (`empirical-inquiry`). One command is cheap; inheriting a false premise is not.
21
+ - **Whoever clears the condition deletes the claim.** Removing the block is not done until the note describing it is gone. A cleared gate whose comment survives is the next reader's false premise, and you are the last person who knows it is false.
22
+ - **Prefer expiry-resistant forms.** State what **is** true rather than what is pending. Where a temporal claim is unavoidable, anchor it to something that fails when it goes stale — a check, a linked work item — never to prose nobody re-reads.
23
+ - **A prediction on a closing item becomes tracked work or is retracted.** If the warning is real it is a work item with an explicit blocking link (`tracked-work`); if it is not real, retract it. Prose on a closed item is neither, and the one-way lifecycle means only `claim-archaeology` can recover it.
24
+ - **A gate that can hold work silently is a defect in the gate.** Any waiting state needs a surfacing mechanism — notification, dashboard, scheduled sweep. Fix the gate; do not resolve to look more often.
25
+
26
+ Full prose, worked examples, and the rewrite patterns: [reference/stale-state-claims.md](../reference/stale-state-claims.md).
@@ -89,3 +89,4 @@ Concretely:
89
89
  - **`verification`** — proves the software behaves as a user needs. This rule proves the proof is real. Codified regression tests added under `codify-verification` are subject to this rule: a codified spec that cannot fail is not a regression gate.
90
90
  - **`empirical-inquiry`** — settles an uncertain fact with the cheapest probe. A probe is a check, so it inherits the falsification requirement, including deliberate bite controls (cases that MUST report a problem) when the probe's job is to detect problems.
91
91
  - **`claim-evidence-mapping`** — an unfalsified gate cannot back a claim.
92
+ - **`stale-state-claims`** — the sibling failure from the opposite direction: an assertion of state that *could* fail but is never re-evaluated, because nothing re-runs prose. Binding a temporal claim to a check is the recommended fix there, which puts that check under this rule.
@@ -0,0 +1,29 @@
1
+ # Settled Decisions
2
+
3
+ Do not ask the operator to decide something that is already settled by standing preference, evidence you just gathered, or a conventional default. A question is appropriate only when the answer materially changes the work and the answer cannot be derived from available artifacts.
4
+
5
+ ## Decision Sources
6
+
7
+ Check these sources before asking:
8
+
9
+ 1. Standing preferences recorded in memory, project config, project rules, instruction files, or earlier in the same conversation.
10
+ 2. Evidence already gathered during the current run. If research or code inspection resolves the choice, report the decision and cite the evidence.
11
+ 3. Conventional defaults where one option is clearly standard and the alternative needs a reason.
12
+
13
+ If one of those sources answers the question, act on it and state the basis briefly.
14
+
15
+ ## When To Ask
16
+
17
+ Ask only when the answer changes the deliverable and cannot be inferred. That includes product decisions, design vocabulary that does not exist yet, destructive or unsafe assumptions, or choices that would waste substantial work if guessed wrong.
18
+
19
+ When asking, ask once, plainly, with a recommended option.
20
+
21
+ ## Common Violations
22
+
23
+ Half-applying a standing instruction is a violation. If a preference covers a multi-step behavior, apply the whole behavior instead of doing one part and asking about the rest.
24
+
25
+ Asking after the research answered the question is also a violation. When gathered evidence points one way unambiguously, treat it as a decision and make the reasoning auditable in the report.
26
+
27
+ ## Reporting
28
+
29
+ Record settled decisions in a short clause, such as "auto-merge on, per standing preference" or "scoped app-wide, because the Android finding decides it." This gives the operator a chance to correct the decision without requiring an avoidable question first.
@@ -0,0 +1,88 @@
1
+ # Stale State Claims — Reference
2
+
3
+ Eager head: [eager/stale-state-claims.md](../eager/stale-state-claims.md).
4
+
5
+ ## Why this rule exists
6
+
7
+ Most bad documentation is *wrong*. This kind was **right** — and that is what makes it dangerous. A note saying "not wired up yet" earned its credibility honestly, so the next reader has no reason to doubt it. There is no defect to find, no test that goes red, no reviewer who objects. The claim simply keeps being read after the world moved.
8
+
9
+ The cost is not confusion, it is confident misdirection in a specific direction: **toward doing nothing**. An expired blocker never causes someone to break production. It causes work to be skipped, re-planned, or escalated — outcomes that look like caution and are indistinguishable from good judgment at the moment they happen.
10
+
11
+ The `falsifiable-checks` rule covers instruments that cannot fail. This one covers assertions that *could* fail but are never re-evaluated, because nothing re-runs prose. Both produce the same end state — a confident answer with no live evidence behind it — from opposite directions.
12
+
13
+ ## The four failure modes, in detail
14
+
15
+ ### 1. Expired blocker note
16
+
17
+ A comment, docstring, config value, or README line records that something is not available yet. The condition clears. The note does not.
18
+
19
+ Observed: a deployment stage's endpoint was left unset with an adjacent comment explaining that its routing had "not yet been promoted." The routing had been live for days. An agent read the comment, believed it, and planned work to enable a capability that was already enabled — then reported the capability as pending.
20
+
21
+ Countermeasures:
22
+ - Treat any not-yet claim as a **hypothesis with an expiry you cannot see**. Resolve it with one live probe (a request, a query, a status read) before it enters a plan, a status report, or an escalation.
23
+ - Cite the probe, not the comment: "verified live at <time>" is evidence; "the comment says" is hearsay about the past.
24
+ - When you are the one who promotes/enables/provisions the thing, the note describing its absence is part of the change. Grep for it.
25
+
26
+ ### 2. Stale gate marker
27
+
28
+ A config entry, flag, or checklist item is annotated as awaiting a human decision or an unfulfilled prerequisite. The prerequisite arrives, or the decision gets made elsewhere. The marker persists and keeps routing work to a gate that is no longer closed.
29
+
30
+ Observed: an entry marked human-gated because no client had been provisioned. The client had existed for twelve days. The gate was honored anyway and the question escalated to a person who had already decided it and was surprised to be asked.
31
+
32
+ This mode is corrosive twice over: it wastes the human's attention, and it teaches the agent that gates are noise — which is exactly the wrong lesson to carry into a gate that is still real.
33
+
34
+ Countermeasures:
35
+ - Before honoring a gate, verify the condition that justifies it still holds. A gate whose stated reason is falsifiable and false is not a gate.
36
+ - Record gates as **conditions**, not as verdicts: "gated until <checkable condition>" can be evaluated; "human-gated" cannot.
37
+ - When escalating, state the evidence that the gate is still live. If you cannot, you are escalating a comment.
38
+
39
+ ### 3. Prediction buried by closure
40
+
41
+ Someone records a correct warning — "this will fail until X is fixed", "this leaves Y broken" — as a comment on a work item, and the item then closes. Closure is a filter: closed items are not read. The prediction is deleted in every practical sense while feeling like it was recorded.
42
+
43
+ Observed: a comment correctly predicted a defect. The item closed 25 minutes later with no follow-up filed. The defect sat unnoticed for **eight days**, blocking a dependent item that was itself parked waiting for it — two items, both stalled, and the explanation was already written down in a place nobody would look.
44
+
45
+ Countermeasures:
46
+ - A prediction is either **real work or a retraction**. If real, it is a tracked leaf with an explicit blocking link to what it blocks (`tracked-work`), created **before** the parent item closes. If not real, say so in the same thread so the next reader is not left holding an unresolved warning.
47
+ - Never let closing be the last action on an item that carries an open prediction. Check the comment thread as part of closing.
48
+ - Blocking links are the mechanism that makes the parked dependent item legible. A prediction with no link is invisible from the side that is actually waiting.
49
+
50
+ ### 4. Silent waiting gate
51
+
52
+ An approval, queue, or review state holds work and emits nothing. Its duration is therefore set by how often somebody happens to look, which is not a property of the work's urgency.
53
+
54
+ Observed: a deploy approval gate held a security-relevant fix undeployed for two days. Nothing was broken, nobody was wrong, and nothing surfaced that the gate was waiting.
55
+
56
+ Countermeasures:
57
+ - Every waiting state needs a surfacing mechanism proportional to what it can hold: notification, dashboard row, scheduled sweep, or an automatic expiry.
58
+ - **Fix the gate, not your habits.** "Remember to check" is not a mechanism; it is the same failure with a person's name on it.
59
+ - When you add a gate, state how long it may silently hold work and what surfaces it. If the answer is "indefinitely" and "nothing", the gate is not finished.
60
+
61
+ ## Expiry-resistant forms
62
+
63
+ Prefer, in order:
64
+
65
+ 1. **State what is true.** "Serves from <path>" outlives "not yet serving from <path>", because it stays wrong-detectably wrong instead of quietly-stale.
66
+ 2. **Bind the claim to a check.** A test, guard, or assertion that fails when the temporal claim goes false converts a silent expiry into a red build. (That check is itself subject to `falsifiable-checks` — a guard that cannot fail re-creates the problem it was added to solve.)
67
+ 3. **Bind the claim to a work item.** "Blocked by <ref>" is resolvable by anyone; "waiting on the migration" is resolvable only by whoever wrote it.
68
+ 4. **Timestamp and scope it.** If prose is genuinely the only option, write "as of <date>" and name the condition that ends it. A dated claim at least advertises its own age.
69
+
70
+ Avoid: bare "not yet", "TODO once", "temporarily", "for now", "pending" with no owner, condition, date, or link. Each is a claim that can only be falsified by someone who already knows the truth — which is the one person who does not need to read it.
71
+
72
+ ## How to apply
73
+
74
+ 1. Reading a note that asserts a pending/blocked/not-yet state: **do not act on it yet.**
75
+ 2. Identify the cheapest live probe that settles the underlying fact.
76
+ 3. Run it. Report the observation, not the note.
77
+ 4. If the note is stale, **delete or correct it in the same change** — leaving it is handing the next reader the trap you just escaped.
78
+ 5. If the note is accurate, upgrade it while you are there: attach the condition, the link, or the check that will make it self-expiring.
79
+
80
+ And when you are the one clearing a condition: sweep for the notes that described it. The change is not complete while the repository still asserts the old state.
81
+
82
+ ## Interaction with other rules
83
+
84
+ - **`falsifiable-checks`** — the sibling failure. That rule prevents an instrument that cannot fail; this one prevents an assertion that never gets re-evaluated. A check bound to a temporal claim (form 2 above) sits under both rules at once.
85
+ - **`empirical-inquiry`** — supplies the discipline this rule depends on: settle the fact with the cheapest probe rather than reasoning from what is written down. A stale note is exactly the "confident-sounding answer from a prior assumption" that rule forbids.
86
+ - **`verification`** — a status derived from a comment is not runtime evidence. Reporting "still blocked" on the strength of a note is the same error as reporting "works" on the strength of the code looking correct.
87
+ - **`tracked-work`** — the destination for any prediction that survives its item's closure: one live leaf, explicit blocking link, carried on the branch and PR.
88
+ - **`claim-archaeology`** — the recovery path once mode 3 has already happened. Lifecycles are one-way, so a warning lost to closure resurfaces only as a fresh item whose ancestry has to be reconstructed. Archaeology is the cleanup; this rule is the prevention.
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "lisa",
3
- "version": "2.314.0",
3
+ "version": "2.316.0",
4
4
  "description": "Universal governance — agents, skills, commands, hooks, and rules for all projects",
5
5
  "author": {
6
6
  "name": "Cody Swann"
@@ -132,6 +132,7 @@ STATE_DIR="${TMPDIR:-/tmp}/lisa-verification-gate"
132
132
  mkdir -p "$STATE_DIR" 2>/dev/null || exit 0
133
133
 
134
134
  ARM_FLAG="${STATE_DIR}/${SESSION_ID}.armed"
135
+ ARM_PLAN_FILE="${STATE_DIR}/${SESSION_ID}.plan"
135
136
  SUBAGENT_FLAG="${STATE_DIR}/${SESSION_ID}.subagent"
136
137
  COUNT_FILE="${STATE_DIR}/${SESSION_ID}.blocks"
137
138
 
@@ -149,6 +150,19 @@ arm_once() {
149
150
  [ -f "$ARM_FLAG" ] || touch "$ARM_FLAG" 2>/dev/null || true
150
151
  }
151
152
 
153
+ record_current_plan() {
154
+ local plan="$1"
155
+ [ -n "$plan" ] || return 0
156
+ [ -f "$ARM_PLAN_FILE" ] && return 0
157
+ printf '%s\n' "$plan" > "$ARM_PLAN_FILE" 2>/dev/null || true
158
+ }
159
+
160
+ plan_from_prompt() {
161
+ printf '%s' "$1" |
162
+ sed -nE '1{s/^[[:space:]]*\/(lisa:implement|implement)[[:space:]]+([^[:space:]]+).*$/\2/p;}' |
163
+ tr '[:upper:]' '[:lower:]'
164
+ }
165
+
152
166
  case "$HOOK_EVENT" in
153
167
  SubagentStart)
154
168
  touch "$SUBAGENT_FLAG" 2>/dev/null || true
@@ -162,6 +176,7 @@ case "$HOOK_EVENT" in
162
176
  case "$LEADING" in
163
177
  /lisa:implement*|/implement*)
164
178
  arm_once
179
+ record_current_plan "$(plan_from_prompt "$LEADING")"
165
180
  ;;
166
181
  esac
167
182
  fi
@@ -175,6 +190,8 @@ case "$HOOK_EVENT" in
175
190
  case "$SKILL_NAME" in
176
191
  lisa-implement|implement)
177
192
  arm_once
193
+ SKILL_ARGUMENTS=$(printf '%s' "$INPUT" | jq -r '.tool_input.arguments // .tool_input.input // empty' 2>/dev/null || true)
194
+ record_current_plan "$(printf '%s' "$SKILL_ARGUMENTS" | awk '{print tolower($1)}')"
178
195
  ;;
179
196
  esac
180
197
  fi
@@ -205,6 +222,39 @@ fi
205
222
  PROJECT_DIR="${CLAUDE_PROJECT_DIR:-.}"
206
223
  VERDICT_FILE="${PROJECT_DIR}/.lisa/verification-status.json"
207
224
 
225
+ # A completed flow's verdict is a shipped record. The next flow in the same
226
+ # worktree writes the SAME path and destroys it — losing the evidence that
227
+ # proved the earlier work, and gating the new run against a verdict whose
228
+ # `plan` names something else entirely. Preserve any verdict belonging to a
229
+ # different plan before this run can overwrite it.
230
+ #
231
+ # Keyed on `.plan`, so re-running the same plan still overwrites in place and
232
+ # no archive accumulates.
233
+ preserve_foreign_verdict() {
234
+ [ -f "$VERDICT_FILE" ] || return 0
235
+ command -v jq >/dev/null 2>&1 || return 0
236
+
237
+ local prior_plan current_plan archive
238
+ prior_plan=$(jq -r '.plan // empty' "$VERDICT_FILE" 2>/dev/null || true)
239
+ [ -n "$prior_plan" ] || return 0
240
+ case "$prior_plan" in
241
+ *[!A-Za-z0-9._-]*)
242
+ return 0
243
+ ;;
244
+ esac
245
+ current_plan=$(cat "$ARM_PLAN_FILE" 2>/dev/null || true)
246
+ [ -z "$current_plan" ] || [ "$prior_plan" != "$current_plan" ] || return 0
247
+
248
+ # Only archive a verdict written BEFORE this flow armed — anything newer
249
+ # belongs to the current run.
250
+ [ "$VERDICT_FILE" -ot "$ARM_FLAG" ] || return 0
251
+
252
+ archive="${PROJECT_DIR}/.lisa/verification-status.${prior_plan}.json"
253
+ [ -f "$archive" ] || cp -p "$VERDICT_FILE" "$archive" 2>/dev/null || true
254
+ }
255
+
256
+ preserve_foreign_verdict
257
+
208
258
  # Set by the v2 path when a claim/evidence violation is what closed the gate,
209
259
  # so the block message can state the real reason instead of the v1 fallback.
210
260
  V2_BLOCK_REASON=""
@@ -387,7 +437,7 @@ verdict_is_terminal() {
387
437
  if verdict_is_terminal; then
388
438
  # Gate satisfied — disarm so a follow-up stop in the same session is not
389
439
  # re-gated against the now-consumed verdict, and allow the stop.
390
- rm -f "$ARM_FLAG" "$COUNT_FILE" 2>/dev/null || true
440
+ rm -f "$ARM_FLAG" "$ARM_PLAN_FILE" "$COUNT_FILE" 2>/dev/null || true
391
441
  exit 0
392
442
  fi
393
443
 
@@ -400,7 +450,7 @@ COUNT=$((COUNT + 1))
400
450
  echo "$COUNT" > "$COUNT_FILE" 2>/dev/null || true
401
451
 
402
452
  if [ "$COUNT" -gt "$MAX_BLOCKS" ]; then
403
- rm -f "$ARM_FLAG" "$COUNT_FILE" 2>/dev/null || true
453
+ rm -f "$ARM_FLAG" "$ARM_PLAN_FILE" "$COUNT_FILE" 2>/dev/null || true
404
454
  cat >&2 <<EOF
405
455
  Verification gate: still no passing verdict after ${MAX_BLOCKS} attempts.
406
456
  Releasing the stop gate to avoid an infinite loop. The /lisa:implement Verify
@@ -20,6 +20,31 @@ Membership is **registration, not skill-existence**: a loop is under this contra
20
20
  registered as a scheduled automation, and registering a new one pulls it in automatically. There is
21
21
  no hardcoded roster of loops anywhere.
22
22
 
23
+ ### Interactive flows are members too
24
+
25
+ The outcome vocabulary is **not cron-specific** — it answers "did this need me?", which an operator
26
+ asks of an interactive run exactly as often as of a scheduled one. Every Lisa flow that terminates
27
+ is a member: `lisa-implement`, `lisa-verify`, `lisa-plan`, `lisa-git-submit-pr`,
28
+ `lisa-drive-pr-to-merge`, `lisa-research`, and any skill invoked as a slash command.
29
+
30
+ For an interactive flow the required run record is the **final user-facing message**, not a JSONL
31
+ row written by `automation-run-record.mjs`. Interactive flows may also record local JSONL telemetry
32
+ when a specific skill owns that surface, but this contract's mandatory record is the final answer.
33
+ It opens with the outcome and the operator action, before any narrative:
34
+
35
+ ```
36
+ change-proved — nothing for you. PR #6393 open, auto-merge on, CI running.
37
+ approval-requested — need a decision: ship the 135 Regular→Bold flips, or hold for design sign-off?
38
+ recovery-required — need you: staging E2E gate red for congestion; rerun or admin-merge.
39
+ ```
40
+
41
+ The operator must learn whether they are needed **from the first line**, without reading the report.
42
+ Findings, evidence and caveats follow; they never replace the action line and never precede it.
43
+
44
+ An interactive flow that ends in prose with the action buried — or absent — is the same contract
45
+ violation as a silent cron exit. "I flagged X, I noticed Y, worth knowing Z" is narrative, not an
46
+ outcome. If nothing is needed, say **"nothing for you"** in those words and stop.
47
+
23
48
  ## The six run outcomes
24
49
 
25
50
  Exactly one per run:
@@ -21,6 +21,7 @@ Do not reason your way to a confident-sounding answer from documentation, prior
21
21
  - Presenting a guess, recollection, or doc summary as established fact when it was cheap to verify and you did not.
22
22
  - "Should work" / "probably" / "the docs say" as the basis for a load-bearing decision an experiment could have settled.
23
23
  - Skipping the probe because the answer "seems obvious" — those are exactly the ones that quietly drift from reality.
24
+ - Treating a recorded "not yet" / "pending" / "blocked" / "human-gated" note as current state. It is a claim about the day it was written; probe the live state before planning around it, escalating it, or reporting it as a blocker (`stale-state-claims`).
24
25
 
25
26
  This is the inquiry counterpart to the `verification` rule (which proves completed work behaves correctly). Both reject "it looks correct" as evidence.
26
27
 
@@ -94,3 +94,4 @@ Concretely:
94
94
  - **`verification`** — proves the software behaves as a user needs. This rule proves the proof is real. Codified regression tests added under `codify-verification` are subject to this rule: a codified spec that cannot fail is not a regression gate.
95
95
  - **`empirical-inquiry`** — settles an uncertain fact with the cheapest probe. A probe is a check, so it inherits the falsification requirement, including deliberate bite controls (cases that MUST report a problem) when the probe's job is to detect problems.
96
96
  - **`claim-evidence-mapping`** — an unfalsified gate cannot back a claim.
97
+ - **`stale-state-claims`** — the sibling failure from the opposite direction: an assertion of state that *could* fail but is never re-evaluated, because nothing re-runs prose. Binding a temporal claim to a check is the recommended fix there, which puts that check under this rule.
@@ -0,0 +1,34 @@
1
+ ---
2
+ description: "Settled Decisions"
3
+ alwaysApply: false
4
+ ---
5
+
6
+ # Settled Decisions
7
+
8
+ Do not ask the operator to decide something that is already settled by standing preference, evidence you just gathered, or a conventional default. A question is appropriate only when the answer materially changes the work and the answer cannot be derived from available artifacts.
9
+
10
+ ## Decision Sources
11
+
12
+ Check these sources before asking:
13
+
14
+ 1. Standing preferences recorded in memory, project config, project rules, instruction files, or earlier in the same conversation.
15
+ 2. Evidence already gathered during the current run. If research or code inspection resolves the choice, report the decision and cite the evidence.
16
+ 3. Conventional defaults where one option is clearly standard and the alternative needs a reason.
17
+
18
+ If one of those sources answers the question, act on it and state the basis briefly.
19
+
20
+ ## When To Ask
21
+
22
+ Ask only when the answer changes the deliverable and cannot be inferred. That includes product decisions, design vocabulary that does not exist yet, destructive or unsafe assumptions, or choices that would waste substantial work if guessed wrong.
23
+
24
+ When asking, ask once, plainly, with a recommended option.
25
+
26
+ ## Common Violations
27
+
28
+ Half-applying a standing instruction is a violation. If a preference covers a multi-step behavior, apply the whole behavior instead of doing one part and asking about the rest.
29
+
30
+ Asking after the research answered the question is also a violation. When gathered evidence points one way unambiguously, treat it as a decision and make the reasoning auditable in the report.
31
+
32
+ ## Reporting
33
+
34
+ Record settled decisions in a short clause, such as "auto-merge on, per standing preference" or "scoped app-wide, because the Android finding decides it." This gives the operator a chance to correct the decision without requiring an avoidable question first.
@@ -0,0 +1,60 @@
1
+ ---
2
+ description: "Settled Decisions (load-bearing)"
3
+ alwaysApply: true
4
+ ---
5
+
6
+ # Settled Decisions (load-bearing)
7
+
8
+ **Never ask the operator a question that a standing preference or your own gathered evidence has
9
+ already answered.** Re-asking a settled decision is not caution — it hands back a choice the
10
+ operator already made and makes them make it twice.
11
+
12
+ ## The test
13
+
14
+ Before asking anything, check the three sources that may already hold the answer:
15
+
16
+ 1. **A standing preference** — something the operator has told you once and expects to hold:
17
+ recorded in project memory, `.lisa.config.json`, a project rule, `CLAUDE.md`/`AGENTS.md`, or
18
+ stated earlier in this conversation. "Every PR gets auto-merge and gets watched to merge" is a
19
+ standing preference; asking "want me to watch this PR?" violates it.
20
+ 2. **Evidence you already gathered** — if the research, measurement or code read you just performed
21
+ resolves the question, the question is answered. Report the decision and the evidence for it.
22
+ 3. **A convention with an obvious default** — where one option is clearly conventional and the other
23
+ needs a reason, take the conventional one and say so in one line.
24
+
25
+ If any source answers it: **act, and state the decision in a clause.** Do not convert it into a
26
+ question.
27
+
28
+ ## When asking IS right
29
+
30
+ Ask when the answer would change the work *and* you genuinely cannot derive it:
31
+
32
+ - The options lead to materially different deliverables and nothing in scope decides between them.
33
+ - Proceeding on a wrong assumption would be unsafe, destructive, or waste substantial work.
34
+ - The answer is a human judgement the artifacts do not contain — a product decision, a design
35
+ vocabulary that does not exist yet, a risk the operator owns.
36
+
37
+ That kind of question is load-bearing and should be asked plainly, once, with a recommendation.
38
+
39
+ ## The failure mode this rule exists to stop
40
+
41
+ Asking feels collaborative, so it gets over-applied — especially at the end of a report, where a
42
+ trailing question reads as deference. It is not deference when the answer was already given; it is
43
+ work handed back. Two specific tells:
44
+
45
+ - **Half-applying an instruction.** Doing the first half of a standing preference automatically and
46
+ asking permission for the second half of the same preference.
47
+ - **Asking after the research answered it.** Completing an investigation that points one direction
48
+ unambiguously, then presenting the conclusion as an open question.
49
+
50
+ ## Partial application is worse than either extreme
51
+
52
+ If a standing preference covers a multi-step behavior, apply **all** of it. Applying part and asking
53
+ about the rest produces the worst outcome: the operator is interrupted *and* the instruction was not
54
+ honored.
55
+
56
+ ## Recording, not asking
57
+
58
+ When you take a settled decision, make it auditable in one clause — "auto-merge on, per your standing
59
+ preference", "scoped app-wide, because the Android finding decides it". That gives the operator the
60
+ chance to correct it without requiring them to answer first.
@@ -0,0 +1,93 @@
1
+ ---
2
+ description: "Stale State Claims — Reference"
3
+ alwaysApply: false
4
+ ---
5
+
6
+ # Stale State Claims — Reference
7
+
8
+ Eager head: [eager/stale-state-claims.md](stale-state-claims.mdc).
9
+
10
+ ## Why this rule exists
11
+
12
+ Most bad documentation is *wrong*. This kind was **right** — and that is what makes it dangerous. A note saying "not wired up yet" earned its credibility honestly, so the next reader has no reason to doubt it. There is no defect to find, no test that goes red, no reviewer who objects. The claim simply keeps being read after the world moved.
13
+
14
+ The cost is not confusion, it is confident misdirection in a specific direction: **toward doing nothing**. An expired blocker never causes someone to break production. It causes work to be skipped, re-planned, or escalated — outcomes that look like caution and are indistinguishable from good judgment at the moment they happen.
15
+
16
+ The `falsifiable-checks` rule covers instruments that cannot fail. This one covers assertions that *could* fail but are never re-evaluated, because nothing re-runs prose. Both produce the same end state — a confident answer with no live evidence behind it — from opposite directions.
17
+
18
+ ## The four failure modes, in detail
19
+
20
+ ### 1. Expired blocker note
21
+
22
+ A comment, docstring, config value, or README line records that something is not available yet. The condition clears. The note does not.
23
+
24
+ Observed: a deployment stage's endpoint was left unset with an adjacent comment explaining that its routing had "not yet been promoted." The routing had been live for days. An agent read the comment, believed it, and planned work to enable a capability that was already enabled — then reported the capability as pending.
25
+
26
+ Countermeasures:
27
+ - Treat any not-yet claim as a **hypothesis with an expiry you cannot see**. Resolve it with one live probe (a request, a query, a status read) before it enters a plan, a status report, or an escalation.
28
+ - Cite the probe, not the comment: "verified live at <time>" is evidence; "the comment says" is hearsay about the past.
29
+ - When you are the one who promotes/enables/provisions the thing, the note describing its absence is part of the change. Grep for it.
30
+
31
+ ### 2. Stale gate marker
32
+
33
+ A config entry, flag, or checklist item is annotated as awaiting a human decision or an unfulfilled prerequisite. The prerequisite arrives, or the decision gets made elsewhere. The marker persists and keeps routing work to a gate that is no longer closed.
34
+
35
+ Observed: an entry marked human-gated because no client had been provisioned. The client had existed for twelve days. The gate was honored anyway and the question escalated to a person who had already decided it and was surprised to be asked.
36
+
37
+ This mode is corrosive twice over: it wastes the human's attention, and it teaches the agent that gates are noise — which is exactly the wrong lesson to carry into a gate that is still real.
38
+
39
+ Countermeasures:
40
+ - Before honoring a gate, verify the condition that justifies it still holds. A gate whose stated reason is falsifiable and false is not a gate.
41
+ - Record gates as **conditions**, not as verdicts: "gated until <checkable condition>" can be evaluated; "human-gated" cannot.
42
+ - When escalating, state the evidence that the gate is still live. If you cannot, you are escalating a comment.
43
+
44
+ ### 3. Prediction buried by closure
45
+
46
+ Someone records a correct warning — "this will fail until X is fixed", "this leaves Y broken" — as a comment on a work item, and the item then closes. Closure is a filter: closed items are not read. The prediction is deleted in every practical sense while feeling like it was recorded.
47
+
48
+ Observed: a comment correctly predicted a defect. The item closed 25 minutes later with no follow-up filed. The defect sat unnoticed for **eight days**, blocking a dependent item that was itself parked waiting for it — two items, both stalled, and the explanation was already written down in a place nobody would look.
49
+
50
+ Countermeasures:
51
+ - A prediction is either **real work or a retraction**. If real, it is a tracked leaf with an explicit blocking link to what it blocks (`tracked-work`), created **before** the parent item closes. If not real, say so in the same thread so the next reader is not left holding an unresolved warning.
52
+ - Never let closing be the last action on an item that carries an open prediction. Check the comment thread as part of closing.
53
+ - Blocking links are the mechanism that makes the parked dependent item legible. A prediction with no link is invisible from the side that is actually waiting.
54
+
55
+ ### 4. Silent waiting gate
56
+
57
+ An approval, queue, or review state holds work and emits nothing. Its duration is therefore set by how often somebody happens to look, which is not a property of the work's urgency.
58
+
59
+ Observed: a deploy approval gate held a security-relevant fix undeployed for two days. Nothing was broken, nobody was wrong, and nothing surfaced that the gate was waiting.
60
+
61
+ Countermeasures:
62
+ - Every waiting state needs a surfacing mechanism proportional to what it can hold: notification, dashboard row, scheduled sweep, or an automatic expiry.
63
+ - **Fix the gate, not your habits.** "Remember to check" is not a mechanism; it is the same failure with a person's name on it.
64
+ - When you add a gate, state how long it may silently hold work and what surfaces it. If the answer is "indefinitely" and "nothing", the gate is not finished.
65
+
66
+ ## Expiry-resistant forms
67
+
68
+ Prefer, in order:
69
+
70
+ 1. **State what is true.** "Serves from <path>" outlives "not yet serving from <path>", because it stays wrong-detectably wrong instead of quietly-stale.
71
+ 2. **Bind the claim to a check.** A test, guard, or assertion that fails when the temporal claim goes false converts a silent expiry into a red build. (That check is itself subject to `falsifiable-checks` — a guard that cannot fail re-creates the problem it was added to solve.)
72
+ 3. **Bind the claim to a work item.** "Blocked by <ref>" is resolvable by anyone; "waiting on the migration" is resolvable only by whoever wrote it.
73
+ 4. **Timestamp and scope it.** If prose is genuinely the only option, write "as of <date>" and name the condition that ends it. A dated claim at least advertises its own age.
74
+
75
+ Avoid: bare "not yet", "TODO once", "temporarily", "for now", "pending" with no owner, condition, date, or link. Each is a claim that can only be falsified by someone who already knows the truth — which is the one person who does not need to read it.
76
+
77
+ ## How to apply
78
+
79
+ 1. Reading a note that asserts a pending/blocked/not-yet state: **do not act on it yet.**
80
+ 2. Identify the cheapest live probe that settles the underlying fact.
81
+ 3. Run it. Report the observation, not the note.
82
+ 4. If the note is stale, **delete or correct it in the same change** — leaving it is handing the next reader the trap you just escaped.
83
+ 5. If the note is accurate, upgrade it while you are there: attach the condition, the link, or the check that will make it self-expiring.
84
+
85
+ And when you are the one clearing a condition: sweep for the notes that described it. The change is not complete while the repository still asserts the old state.
86
+
87
+ ## Interaction with other rules
88
+
89
+ - **`falsifiable-checks`** — the sibling failure. That rule prevents an instrument that cannot fail; this one prevents an assertion that never gets re-evaluated. A check bound to a temporal claim (form 2 above) sits under both rules at once.
90
+ - **`empirical-inquiry`** — supplies the discipline this rule depends on: settle the fact with the cheapest probe rather than reasoning from what is written down. A stale note is exactly the "confident-sounding answer from a prior assumption" that rule forbids.
91
+ - **`verification`** — a status derived from a comment is not runtime evidence. Reporting "still blocked" on the strength of a note is the same error as reporting "works" on the strength of the code looking correct.
92
+ - **`tracked-work`** — the destination for any prediction that survives its item's closure: one live leaf, explicit blocking link, carried on the branch and PR.
93
+ - **`claim-archaeology`** — the recovery path once mode 3 has already happened. Lifecycles are one-way, so a warning lost to closure resurfaces only as a fresh item whose ancestry has to be reconstructed. Archaeology is the cleanup; this rule is the prevention.
@@ -0,0 +1,31 @@
1
+ ---
2
+ description: "Stale State Claims — \"Not Yet\" Expires, the Note Does Not (load-bearing)"
3
+ alwaysApply: true
4
+ ---
5
+
6
+ # Stale State Claims — "Not Yet" Expires, the Note Does Not (load-bearing)
7
+
8
+ A comment, docstring, config annotation, or work-item note that records a **temporary** state — "not yet", "pending", "waiting on X", "human-gated", "will flip once Y ships" — was accurate the day it was written and is believed long after it stopped being true. Nobody revisits prose when the condition clears.
9
+
10
+ It then misdirects with full authority: work gets skipped as blocked when nothing blocks it, re-planned as undone when it already shipped, or escalated to a person who made that exact decision weeks ago.
11
+
12
+ This is the sibling of `falsifiable-checks`. That rule is about **instruments that cannot fail**; this one is about **assertions of state that have expired**. Both read as authoritative and neither announces its own decay.
13
+
14
+ ## The four ways a recorded state outlives its truth
15
+
16
+ Each has been observed; each was believed long after it went false:
17
+
18
+ 1. **Expired blocker note** — prose asserting a not-yet condition that has since cleared. An endpoint left unset under a comment saying its routing was "not yet promoted" had been serving live for days; an agent planned work to enable what was already enabled.
19
+ 2. **Stale gate marker** — a "pending human approval / not provisioned yet" annotation whose precondition is satisfied. One held for twelve days after the thing it waited on existed, and the decision was escalated to a person surprised to be asked.
20
+ 3. **Prediction buried by closure** — a warning recorded in a comment on an item that then closes. Closure deletes the warning: one correctly predicted a defect, the item closed 25 minutes later, and the defect sat unnoticed for eight days while a dependent item was parked waiting on it.
21
+ 4. **Silent waiting gate** — a queue or approval state with no surfacing mechanism, which holds work for exactly as long as nobody happens to look. One held a security-relevant fix undeployed for two days.
22
+
23
+ ## Mandatory
24
+
25
+ - **A recorded blocker is a claim about the past — check the present before acting on it.** Before planning around it, escalating it, or reporting it as a blocker, probe the live state (`empirical-inquiry`). One command is cheap; inheriting a false premise is not.
26
+ - **Whoever clears the condition deletes the claim.** Removing the block is not done until the note describing it is gone. A cleared gate whose comment survives is the next reader's false premise, and you are the last person who knows it is false.
27
+ - **Prefer expiry-resistant forms.** State what **is** true rather than what is pending. Where a temporal claim is unavoidable, anchor it to something that fails when it goes stale — a check, a linked work item — never to prose nobody re-reads.
28
+ - **A prediction on a closing item becomes tracked work or is retracted.** If the warning is real it is a work item with an explicit blocking link (`tracked-work`); if it is not real, retract it. Prose on a closed item is neither, and the one-way lifecycle means only `claim-archaeology` can recover it.
29
+ - **A gate that can hold work silently is a defect in the gate.** Any waiting state needs a surfacing mechanism — notification, dashboard, scheduled sweep. Fix the gate; do not resolve to look more often.
30
+
31
+ Full prose, worked examples, and the rewrite patterns: [reference/stale-state-claims.md](stale-state-claims-reference.mdc).
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "lisa-expo",
3
- "version": "2.314.0",
3
+ "version": "2.316.0",
4
4
  "description": "Expo/React Native-specific skills, agents, rules, and MCP servers",
5
5
  "author": {
6
6
  "name": "Cody Swann"
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "lisa-expo",
3
- "version": "2.314.0",
3
+ "version": "2.316.0",
4
4
  "description": "Expo and React Native-specific skills, agents, rules, and MCP servers.",
5
5
  "author": {
6
6
  "name": "Cody Swann"
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "lisa-expo",
3
- "version": "2.314.0",
3
+ "version": "2.316.0",
4
4
  "description": "Expo/React Native-specific skills, agents, rules, and MCP servers",
5
5
  "author": {
6
6
  "name": "Cody Swann"