@codyswann/lisa 2.313.2 → 2.315.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/core/upstream-evidence-manifest.d.ts.map +1 -1
- package/dist/core/upstream-evidence-manifest.js +27 -5
- package/dist/core/upstream-evidence-manifest.js.map +1 -1
- package/package.json +1 -1
- package/plugins/lisa/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa/.codex-plugin/skills/lisa-codify-verification/SKILL.md +23 -7
- package/plugins/lisa/agents/test-specialist.md +3 -1
- package/plugins/lisa/agents/verification-specialist.md +4 -1
- package/plugins/lisa/rules/eager/empirical-inquiry.md +1 -0
- package/plugins/lisa/rules/eager/falsifiable-checks.md +26 -0
- package/plugins/lisa/rules/eager/stale-state-claims.md +26 -0
- package/plugins/lisa/rules/eager/verification.md +1 -0
- package/plugins/lisa/rules/reference/falsifiable-checks.md +92 -0
- package/plugins/lisa/rules/reference/stale-state-claims.md +88 -0
- package/plugins/lisa/skills/lisa-codify-verification/SKILL.md +23 -7
- package/plugins/lisa-agy/agents/test-specialist.md +3 -1
- package/plugins/lisa-agy/agents/verification-specialist.md +4 -1
- package/plugins/lisa-agy/plugin.json +1 -1
- package/plugins/lisa-agy/skills/lisa-codify-verification/SKILL.md +23 -7
- package/plugins/lisa-cdk/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-cdk/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-cdk-agy/plugin.json +1 -1
- package/plugins/lisa-cdk-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-cdk-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-copilot/agents/test-specialist.agent.md +3 -1
- package/plugins/lisa-copilot/agents/verification-specialist.agent.md +4 -1
- package/plugins/lisa-copilot/rules/eager/empirical-inquiry.md +1 -0
- package/plugins/lisa-copilot/rules/eager/falsifiable-checks.md +26 -0
- package/plugins/lisa-copilot/rules/eager/stale-state-claims.md +26 -0
- package/plugins/lisa-copilot/rules/eager/verification.md +1 -0
- package/plugins/lisa-copilot/rules/reference/falsifiable-checks.md +92 -0
- package/plugins/lisa-copilot/rules/reference/stale-state-claims.md +88 -0
- package/plugins/lisa-copilot/skills/lisa-codify-verification/SKILL.md +23 -7
- package/plugins/lisa-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-cursor/agents/test-specialist.md +3 -1
- package/plugins/lisa-cursor/agents/verification-specialist.md +4 -1
- package/plugins/lisa-cursor/rules/empirical-inquiry.mdc +1 -0
- package/plugins/lisa-cursor/rules/falsifiable-checks-reference.mdc +97 -0
- package/plugins/lisa-cursor/rules/falsifiable-checks.mdc +31 -0
- package/plugins/lisa-cursor/rules/stale-state-claims-reference.mdc +93 -0
- package/plugins/lisa-cursor/rules/stale-state-claims.mdc +31 -0
- package/plugins/lisa-cursor/rules/verification.mdc +1 -0
- package/plugins/lisa-cursor/skills/lisa-codify-verification/SKILL.md +23 -7
- package/plugins/lisa-expo/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-expo/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-expo-agy/plugin.json +1 -1
- package/plugins/lisa-expo-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-expo-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-harper-fabric/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-harper-fabric/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-harper-fabric-agy/plugin.json +1 -1
- package/plugins/lisa-harper-fabric-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-harper-fabric-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-nestjs/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-nestjs/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-nestjs-agy/plugin.json +1 -1
- package/plugins/lisa-nestjs-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-nestjs-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-openclaw/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-openclaw/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-openclaw-agy/plugin.json +1 -1
- package/plugins/lisa-openclaw-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-openclaw-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-phaser/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-phaser/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-phaser-agy/plugin.json +1 -1
- package/plugins/lisa-phaser-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-phaser-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-rails/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-rails/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-rails-agy/plugin.json +1 -1
- package/plugins/lisa-rails-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-rails-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-typescript/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-typescript/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-typescript-agy/plugin.json +1 -1
- package/plugins/lisa-typescript-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-typescript-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-wiki/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-wiki/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-wiki-agy/plugin.json +1 -1
- package/plugins/lisa-wiki-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-wiki-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/src/base/agents/test-specialist.md +3 -1
- package/plugins/src/base/agents/verification-specialist.md +4 -1
- package/plugins/src/base/rules/eager/empirical-inquiry.md +1 -0
- package/plugins/src/base/rules/eager/falsifiable-checks.md +26 -0
- package/plugins/src/base/rules/eager/stale-state-claims.md +26 -0
- package/plugins/src/base/rules/eager/verification.md +1 -0
- package/plugins/src/base/rules/reference/falsifiable-checks.md +92 -0
- package/plugins/src/base/rules/reference/stale-state-claims.md +88 -0
- package/plugins/src/base/skills/lisa-codify-verification/SKILL.md +23 -7
|
@@ -0,0 +1,93 @@
|
|
|
1
|
+
---
|
|
2
|
+
description: "Stale State Claims — Reference"
|
|
3
|
+
alwaysApply: false
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Stale State Claims — Reference
|
|
7
|
+
|
|
8
|
+
Eager head: [eager/stale-state-claims.md](stale-state-claims.mdc).
|
|
9
|
+
|
|
10
|
+
## Why this rule exists
|
|
11
|
+
|
|
12
|
+
Most bad documentation is *wrong*. This kind was **right** — and that is what makes it dangerous. A note saying "not wired up yet" earned its credibility honestly, so the next reader has no reason to doubt it. There is no defect to find, no test that goes red, no reviewer who objects. The claim simply keeps being read after the world moved.
|
|
13
|
+
|
|
14
|
+
The cost is not confusion, it is confident misdirection in a specific direction: **toward doing nothing**. An expired blocker never causes someone to break production. It causes work to be skipped, re-planned, or escalated — outcomes that look like caution and are indistinguishable from good judgment at the moment they happen.
|
|
15
|
+
|
|
16
|
+
The `falsifiable-checks` rule covers instruments that cannot fail. This one covers assertions that *could* fail but are never re-evaluated, because nothing re-runs prose. Both produce the same end state — a confident answer with no live evidence behind it — from opposite directions.
|
|
17
|
+
|
|
18
|
+
## The four failure modes, in detail
|
|
19
|
+
|
|
20
|
+
### 1. Expired blocker note
|
|
21
|
+
|
|
22
|
+
A comment, docstring, config value, or README line records that something is not available yet. The condition clears. The note does not.
|
|
23
|
+
|
|
24
|
+
Observed: a deployment stage's endpoint was left unset with an adjacent comment explaining that its routing had "not yet been promoted." The routing had been live for days. An agent read the comment, believed it, and planned work to enable a capability that was already enabled — then reported the capability as pending.
|
|
25
|
+
|
|
26
|
+
Countermeasures:
|
|
27
|
+
- Treat any not-yet claim as a **hypothesis with an expiry you cannot see**. Resolve it with one live probe (a request, a query, a status read) before it enters a plan, a status report, or an escalation.
|
|
28
|
+
- Cite the probe, not the comment: "verified live at <time>" is evidence; "the comment says" is hearsay about the past.
|
|
29
|
+
- When you are the one who promotes/enables/provisions the thing, the note describing its absence is part of the change. Grep for it.
|
|
30
|
+
|
|
31
|
+
### 2. Stale gate marker
|
|
32
|
+
|
|
33
|
+
A config entry, flag, or checklist item is annotated as awaiting a human decision or an unfulfilled prerequisite. The prerequisite arrives, or the decision gets made elsewhere. The marker persists and keeps routing work to a gate that is no longer closed.
|
|
34
|
+
|
|
35
|
+
Observed: an entry marked human-gated because no client had been provisioned. The client had existed for twelve days. The gate was honored anyway and the question escalated to a person who had already decided it and was surprised to be asked.
|
|
36
|
+
|
|
37
|
+
This mode is corrosive twice over: it wastes the human's attention, and it teaches the agent that gates are noise — which is exactly the wrong lesson to carry into a gate that is still real.
|
|
38
|
+
|
|
39
|
+
Countermeasures:
|
|
40
|
+
- Before honoring a gate, verify the condition that justifies it still holds. A gate whose stated reason is falsifiable and false is not a gate.
|
|
41
|
+
- Record gates as **conditions**, not as verdicts: "gated until <checkable condition>" can be evaluated; "human-gated" cannot.
|
|
42
|
+
- When escalating, state the evidence that the gate is still live. If you cannot, you are escalating a comment.
|
|
43
|
+
|
|
44
|
+
### 3. Prediction buried by closure
|
|
45
|
+
|
|
46
|
+
Someone records a correct warning — "this will fail until X is fixed", "this leaves Y broken" — as a comment on a work item, and the item then closes. Closure is a filter: closed items are not read. The prediction is deleted in every practical sense while feeling like it was recorded.
|
|
47
|
+
|
|
48
|
+
Observed: a comment correctly predicted a defect. The item closed 25 minutes later with no follow-up filed. The defect sat unnoticed for **eight days**, blocking a dependent item that was itself parked waiting for it — two items, both stalled, and the explanation was already written down in a place nobody would look.
|
|
49
|
+
|
|
50
|
+
Countermeasures:
|
|
51
|
+
- A prediction is either **real work or a retraction**. If real, it is a tracked leaf with an explicit blocking link to what it blocks (`tracked-work`), created **before** the parent item closes. If not real, say so in the same thread so the next reader is not left holding an unresolved warning.
|
|
52
|
+
- Never let closing be the last action on an item that carries an open prediction. Check the comment thread as part of closing.
|
|
53
|
+
- Blocking links are the mechanism that makes the parked dependent item legible. A prediction with no link is invisible from the side that is actually waiting.
|
|
54
|
+
|
|
55
|
+
### 4. Silent waiting gate
|
|
56
|
+
|
|
57
|
+
An approval, queue, or review state holds work and emits nothing. Its duration is therefore set by how often somebody happens to look, which is not a property of the work's urgency.
|
|
58
|
+
|
|
59
|
+
Observed: a deploy approval gate held a security-relevant fix undeployed for two days. Nothing was broken, nobody was wrong, and nothing surfaced that the gate was waiting.
|
|
60
|
+
|
|
61
|
+
Countermeasures:
|
|
62
|
+
- Every waiting state needs a surfacing mechanism proportional to what it can hold: notification, dashboard row, scheduled sweep, or an automatic expiry.
|
|
63
|
+
- **Fix the gate, not your habits.** "Remember to check" is not a mechanism; it is the same failure with a person's name on it.
|
|
64
|
+
- When you add a gate, state how long it may silently hold work and what surfaces it. If the answer is "indefinitely" and "nothing", the gate is not finished.
|
|
65
|
+
|
|
66
|
+
## Expiry-resistant forms
|
|
67
|
+
|
|
68
|
+
Prefer, in order:
|
|
69
|
+
|
|
70
|
+
1. **State what is true.** "Serves from <path>" outlives "not yet serving from <path>", because it stays wrong-detectably wrong instead of quietly-stale.
|
|
71
|
+
2. **Bind the claim to a check.** A test, guard, or assertion that fails when the temporal claim goes false converts a silent expiry into a red build. (That check is itself subject to `falsifiable-checks` — a guard that cannot fail re-creates the problem it was added to solve.)
|
|
72
|
+
3. **Bind the claim to a work item.** "Blocked by <ref>" is resolvable by anyone; "waiting on the migration" is resolvable only by whoever wrote it.
|
|
73
|
+
4. **Timestamp and scope it.** If prose is genuinely the only option, write "as of <date>" and name the condition that ends it. A dated claim at least advertises its own age.
|
|
74
|
+
|
|
75
|
+
Avoid: bare "not yet", "TODO once", "temporarily", "for now", "pending" with no owner, condition, date, or link. Each is a claim that can only be falsified by someone who already knows the truth — which is the one person who does not need to read it.
|
|
76
|
+
|
|
77
|
+
## How to apply
|
|
78
|
+
|
|
79
|
+
1. Reading a note that asserts a pending/blocked/not-yet state: **do not act on it yet.**
|
|
80
|
+
2. Identify the cheapest live probe that settles the underlying fact.
|
|
81
|
+
3. Run it. Report the observation, not the note.
|
|
82
|
+
4. If the note is stale, **delete or correct it in the same change** — leaving it is handing the next reader the trap you just escaped.
|
|
83
|
+
5. If the note is accurate, upgrade it while you are there: attach the condition, the link, or the check that will make it self-expiring.
|
|
84
|
+
|
|
85
|
+
And when you are the one clearing a condition: sweep for the notes that described it. The change is not complete while the repository still asserts the old state.
|
|
86
|
+
|
|
87
|
+
## Interaction with other rules
|
|
88
|
+
|
|
89
|
+
- **`falsifiable-checks`** — the sibling failure. That rule prevents an instrument that cannot fail; this one prevents an assertion that never gets re-evaluated. A check bound to a temporal claim (form 2 above) sits under both rules at once.
|
|
90
|
+
- **`empirical-inquiry`** — supplies the discipline this rule depends on: settle the fact with the cheapest probe rather than reasoning from what is written down. A stale note is exactly the "confident-sounding answer from a prior assumption" that rule forbids.
|
|
91
|
+
- **`verification`** — a status derived from a comment is not runtime evidence. Reporting "still blocked" on the strength of a note is the same error as reporting "works" on the strength of the code looking correct.
|
|
92
|
+
- **`tracked-work`** — the destination for any prediction that survives its item's closure: one live leaf, explicit blocking link, carried on the branch and PR.
|
|
93
|
+
- **`claim-archaeology`** — the recovery path once mode 3 has already happened. Lifecycles are one-way, so a warning lost to closure resurfaces only as a fresh item whose ancestry has to be reconstructed. Archaeology is the cleanup; this rule is the prevention.
|
|
@@ -0,0 +1,31 @@
|
|
|
1
|
+
---
|
|
2
|
+
description: "Stale State Claims — \"Not Yet\" Expires, the Note Does Not (load-bearing)"
|
|
3
|
+
alwaysApply: true
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Stale State Claims — "Not Yet" Expires, the Note Does Not (load-bearing)
|
|
7
|
+
|
|
8
|
+
A comment, docstring, config annotation, or work-item note that records a **temporary** state — "not yet", "pending", "waiting on X", "human-gated", "will flip once Y ships" — was accurate the day it was written and is believed long after it stopped being true. Nobody revisits prose when the condition clears.
|
|
9
|
+
|
|
10
|
+
It then misdirects with full authority: work gets skipped as blocked when nothing blocks it, re-planned as undone when it already shipped, or escalated to a person who made that exact decision weeks ago.
|
|
11
|
+
|
|
12
|
+
This is the sibling of `falsifiable-checks`. That rule is about **instruments that cannot fail**; this one is about **assertions of state that have expired**. Both read as authoritative and neither announces its own decay.
|
|
13
|
+
|
|
14
|
+
## The four ways a recorded state outlives its truth
|
|
15
|
+
|
|
16
|
+
Each has been observed; each was believed long after it went false:
|
|
17
|
+
|
|
18
|
+
1. **Expired blocker note** — prose asserting a not-yet condition that has since cleared. An endpoint left unset under a comment saying its routing was "not yet promoted" had been serving live for days; an agent planned work to enable what was already enabled.
|
|
19
|
+
2. **Stale gate marker** — a "pending human approval / not provisioned yet" annotation whose precondition is satisfied. One held for twelve days after the thing it waited on existed, and the decision was escalated to a person surprised to be asked.
|
|
20
|
+
3. **Prediction buried by closure** — a warning recorded in a comment on an item that then closes. Closure deletes the warning: one correctly predicted a defect, the item closed 25 minutes later, and the defect sat unnoticed for eight days while a dependent item was parked waiting on it.
|
|
21
|
+
4. **Silent waiting gate** — a queue or approval state with no surfacing mechanism, which holds work for exactly as long as nobody happens to look. One held a security-relevant fix undeployed for two days.
|
|
22
|
+
|
|
23
|
+
## Mandatory
|
|
24
|
+
|
|
25
|
+
- **A recorded blocker is a claim about the past — check the present before acting on it.** Before planning around it, escalating it, or reporting it as a blocker, probe the live state (`empirical-inquiry`). One command is cheap; inheriting a false premise is not.
|
|
26
|
+
- **Whoever clears the condition deletes the claim.** Removing the block is not done until the note describing it is gone. A cleared gate whose comment survives is the next reader's false premise, and you are the last person who knows it is false.
|
|
27
|
+
- **Prefer expiry-resistant forms.** State what **is** true rather than what is pending. Where a temporal claim is unavoidable, anchor it to something that fails when it goes stale — a check, a linked work item — never to prose nobody re-reads.
|
|
28
|
+
- **A prediction on a closing item becomes tracked work or is retracted.** If the warning is real it is a work item with an explicit blocking link (`tracked-work`); if it is not real, retract it. Prose on a closed item is neither, and the one-way lifecycle means only `claim-archaeology` can recover it.
|
|
29
|
+
- **A gate that can hold work silently is a defect in the gate.** Any waiting state needs a surfacing mechanism — notification, dashboard, scheduled sweep. Fix the gate; do not resolve to look more often.
|
|
30
|
+
|
|
31
|
+
Full prose, worked examples, and the rewrite patterns: [reference/stale-state-claims.md](stale-state-claims-reference.mdc).
|
|
@@ -15,6 +15,7 @@ alwaysApply: true
|
|
|
15
15
|
|
|
16
16
|
- **Never claim success without runtime evidence.** "The code looks correct" is not evidence.
|
|
17
17
|
- **If all you did was run tests, typecheck, and lint — you have NOT verified.**
|
|
18
|
+
- **A check that cannot fail is not evidence either.** Every gate, probe, sweep, and codified spec you author is subject to the `falsifiable-checks` rule: break the guarded property, observe the check fail and name the location, then restore — and report that falsification with the result. "Mentally reverting" does not count, and an unfalsified gate is reported as *unvalidated*, never as passing.
|
|
18
19
|
- **Browser-controller neutrality.** For UI work, control a live browser and perform the Validation Journey as a human would. An in-app Browser/Chrome tool, interactive Playwright control (MCP, API, or ad hoc script), CDP, computer use, the optional Lisa-owned Kane adapter, or an equivalent controller is acceptable. Kane requires explicit upload approval, a passing `lisa kane probe`, an allow-listed non-production environment, and mutation policy `full`; its provider failure is not a product failure. Do not block merely because one preferred backend is unavailable when another interactive controller can drive the browser. Running an automated Playwright or Maestro test alone is still a quality gate, not the initial empirical evidence; after the live journey passes, codify it in the applicable native runner(s). Kane never replaces those regression gates.
|
|
19
20
|
- **Before starting implementation, state your verification plan** — how you will USE the resulting software to prove it works. A plan that only lists `test`/`typecheck`/`lint` commands is not a plan. Do not begin until confirmed.
|
|
20
21
|
- **After verifying empirically, codify it as a regression test** via the `codify-verification` skill — Playwright for UI, integration test for API/DB/auth, benchmark for performance. Codification is mandatory for every verification type except PR/Documentation/Deploy and Investigate-Only spikes. For **frontend work**, codification is dual-runner: a Playwright spec in the project's Playwright test runner AND a Maestro flow in the Maestro test runner whenever the project supports Maestro (`.maestro/` directory, `maestro:test` script, or Maestro CI workflow) — both encoding the same verified journey, neither a substitute for the other.
|
|
@@ -135,9 +135,22 @@ Run only the new test, using whatever per-test invocation the project supports:
|
|
|
135
135
|
|
|
136
136
|
Confirm:
|
|
137
137
|
1. The test PASSES against the current code (the change being shipped)
|
|
138
|
-
2. The test
|
|
138
|
+
2. The test ACTUALLY FAILS without the change — observed, not reasoned about
|
|
139
139
|
|
|
140
|
-
|
|
140
|
+
**Step 2 is mandatory for every codified test, and "mentally reverting" does not satisfy it.** Mental reversion is the exact mechanism by which non-functional guards ship: the author believes the assertion is load-bearing, and it is not. Break the guarded property for real, run the test, and read the failure. See `.claude/rules/falsifiable-checks.md` for the four observed ways a check passes while asserting nothing.
|
|
141
|
+
|
|
142
|
+
Do it one of these ways, in order of preference:
|
|
143
|
+
|
|
144
|
+
- **Run against the pre-fix commit** (bug fixes): check out the failing commit, run the new test, see it fail, return to the fix branch.
|
|
145
|
+
- **Break the property in place**: delete the field, revert the line, flip the condition; run; restore. Prefer this when there is no single pre-fix commit.
|
|
146
|
+
- **Unit-test the checker against synthetic bad input**: required when the input is generated, schema-validated, or cached — a revert can silently fail to change what the test reads (a generator that errors leaves the previous artifact in place, and the test then "passes" on stale input).
|
|
147
|
+
|
|
148
|
+
Two properties the failure itself must have:
|
|
149
|
+
|
|
150
|
+
- It must **name the right location**. A failure that does not localize is weak evidence the test is measuring the intended thing.
|
|
151
|
+
- It must not be satisfiable by the test's own fixture. If the assertion can be met by data the test supplies rather than by the artifact under test, add an explicit assertion against the real artifact (the document, the config, the component's actual output) — or reuse an existing source-bound test.
|
|
152
|
+
|
|
153
|
+
Record the falsification in the codification report (what you broke, how it failed). A codified test whose failure has not been observed is reported as **unvalidated**, not as a regression gate.
|
|
141
154
|
|
|
142
155
|
### 5. Wire it into the suite
|
|
143
156
|
|
|
@@ -162,13 +175,13 @@ Append to the verification report (or PR description):
|
|
|
162
175
|
```markdown
|
|
163
176
|
### Codified Verifications
|
|
164
177
|
|
|
165
|
-
| # | Verification | Framework | Test file | Status |
|
|
166
|
-
|
|
167
|
-
| 1 | <description> | Playwright | `e2e/checkout.spec.ts::displays order confirmation after checkout` | PASS |
|
|
168
|
-
| 2 | <same journey, native surface> | Maestro | `.maestro/flows/checkout-confirmation.yaml` | PASS |
|
|
178
|
+
| # | Verification | Framework | Test file | Status | Falsified by |
|
|
179
|
+
|---|--------------|-----------|-----------|--------|--------------|
|
|
180
|
+
| 1 | <description> | Playwright | `e2e/checkout.spec.ts::displays order confirmation after checkout` | PASS | removed the confirmation render → failed at `checkout.spec.ts:42` |
|
|
181
|
+
| 2 | <same journey, native surface> | Maestro | `.maestro/flows/checkout-confirmation.yaml` | PASS | same break → flow failed on the confirmation assertion |
|
|
169
182
|
```
|
|
170
183
|
|
|
171
|
-
This evidence shows the verification is now guarded.
|
|
184
|
+
This evidence shows the verification is now guarded. **The `Falsified by` column is required** — it names the deliberate break and the observed failure. `UNVALIDATED` is the only permitted alternative, and it means the test is not yet a regression gate.
|
|
172
185
|
|
|
173
186
|
## Output
|
|
174
187
|
|
|
@@ -183,6 +196,9 @@ If codification was skipped, an explicit reason recorded in the report (one of t
|
|
|
183
196
|
## Rules
|
|
184
197
|
|
|
185
198
|
- Never claim a verification is codified without running the new test and observing it pass
|
|
199
|
+
- Never claim it is codified without observing it **FAIL** on a real break — mental reversion is not observation, and a test whose failure was never seen is unvalidated, not a gate (`.claude/rules/falsifiable-checks.md`)
|
|
200
|
+
- Never let the assertion be satisfiable by the test's own fixture instead of the artifact under test — bind it to the real document/config/output
|
|
201
|
+
- Never trust a revert-to-verify on generated, schema-validated, or cached input without confirming the input actually changed; a failed generator silently leaves the old artifact and the test "passes" on stale bytes
|
|
186
202
|
- Never disable, skip, or `.skip()` the new test "temporarily" to make CI green — fix the test or fix the underlying change
|
|
187
203
|
- Never use `expect(true).toBe(true)` placeholders or smoke-only assertions that don't actually exercise the verified behavior
|
|
188
204
|
- Never reuse the verification's manual artifact (screenshot, curl output) as a "test" — those are evidence, not regression coverage
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "lisa-openclaw",
|
|
3
|
-
"version": "2.
|
|
3
|
+
"version": "2.315.0",
|
|
4
4
|
"description": "Connect staff roles to Telegram or Slack via OpenClaw — facilitator/specialist hub-and-spoke routing and repo-coding topics, for Claude Code and Codex",
|
|
5
5
|
"author": {
|
|
6
6
|
"name": "Cody Swann"
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "lisa-openclaw",
|
|
3
|
-
"version": "2.
|
|
3
|
+
"version": "2.315.0",
|
|
4
4
|
"description": "Connect staff roles to Telegram or Slack via OpenClaw — facilitator/specialist hub-and-spoke routing and repo-coding topics, across Claude and Codex.",
|
|
5
5
|
"author": {
|
|
6
6
|
"name": "Cody Swann"
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "lisa-openclaw",
|
|
3
|
-
"version": "2.
|
|
3
|
+
"version": "2.315.0",
|
|
4
4
|
"description": "Connect staff roles to Telegram or Slack via OpenClaw — facilitator/specialist hub-and-spoke routing and repo-coding topics, for Claude Code and Codex",
|
|
5
5
|
"author": {
|
|
6
6
|
"name": "Cody Swann"
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "lisa-openclaw",
|
|
3
|
-
"version": "2.
|
|
3
|
+
"version": "2.315.0",
|
|
4
4
|
"description": "Connect staff roles to Telegram or Slack via OpenClaw — facilitator/specialist hub-and-spoke routing and repo-coding topics, for Claude Code and Codex",
|
|
5
5
|
"author": {
|
|
6
6
|
"name": "Cody Swann"
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "lisa-openclaw",
|
|
3
|
-
"version": "2.
|
|
3
|
+
"version": "2.315.0",
|
|
4
4
|
"description": "Connect staff roles to Telegram or Slack via OpenClaw — facilitator/specialist hub-and-spoke routing and repo-coding topics, for Claude Code and Codex",
|
|
5
5
|
"author": {
|
|
6
6
|
"name": "Cody Swann"
|
|
@@ -21,6 +21,8 @@ You decide what has to be true for this change to be trusted, and design the tes
|
|
|
21
21
|
|
|
22
22
|
Do not write tests against the implementation's shape — they pass through a rewrite that breaks behaviour, which is the opposite of the job. Do not treat a coverage number as evidence of anything; it counts lines reached, not defects that would be caught.
|
|
23
23
|
|
|
24
|
+
Do not hand on a test you have not watched fail. A test that cannot fail is worse than a missing one: it reports the defect as absent and ends the search. Break the behaviour, watch the assertion fail and name the right place, restore. Watch especially for the assertion that is satisfiable by the test's own fixture rather than by the artifact under test — that one passes no matter what the production code does. `.claude/rules/falsifiable-checks.md` has the four observed shapes.
|
|
25
|
+
|
|
24
26
|
## What you hand on
|
|
25
27
|
|
|
26
|
-
The matrix, the edge cases with the reason each is interesting, the TDD sequence, and the commands that run it all. Where behaviour is user-visible, say which runner proves it end to end.
|
|
28
|
+
The matrix, the edge cases with the reason each is interesting, the TDD sequence, and the commands that run it all. Where behaviour is user-visible, say which runner proves it end to end. For each test, what break makes it fail — an assertion whose failure mode you cannot name is not yet designed.
|
|
@@ -12,7 +12,7 @@ skills:
|
|
|
12
12
|
|
|
13
13
|
You are a verification specialist. Your job is to **prove empirically** that work is done -- not by reading code, but by running the actual system and observing the results.
|
|
14
14
|
|
|
15
|
-
Read `.claude/rules/verification.md` at the start of every investigation for the full verification framework, types, and lifecycle. Read `.claude/rules/claim-evidence-mapping.md`
|
|
15
|
+
Read `.claude/rules/verification.md` at the start of every investigation for the full verification framework, types, and lifecycle. Read `.claude/rules/falsifiable-checks.md` alongside it: every check YOU author — probe, script, codified spec, sweep — is subject to it, and a check that has not been shown capable of failing is reported as *unvalidated*, never as passing. Read `.claude/rules/claim-evidence-mapping.md` too: it binds every claim to the **boundary** it asserts and every boundary to the evidence **kinds** that reach it. The verdict you write is what `spec-conformance-specialist` cross-checks — record each claim's `boundary`, its `required_evidence_kinds`, its `evidence_refs`, and its `not_established` list so a boundary mismatch is catchable rather than invisible.
|
|
16
16
|
|
|
17
17
|
## Core Philosophy
|
|
18
18
|
|
|
@@ -128,6 +128,9 @@ For every empirical verification that produced PASS evidence, invoke the `codify
|
|
|
128
128
|
- Follow the verification lifecycle: confirm quality gates, classify, check tooling, fail fast, plan, execute, codify, spec conformance, loop
|
|
129
129
|
- Every passing empirical verification must be codified as a regression test via `codify-verification` before declaring done (skip allowed only for PR / Documentation / Deploy / Investigate-Only)
|
|
130
130
|
- Tests, typecheck, lint, and format are quality gates (prerequisites), NOT verification — never report them as verification evidence
|
|
131
|
+
- Falsify every check you author before reporting a clean result: break the guarded property, confirm the check fails and NAMES the right location, restore. Report what you broke alongside the result — see `.claude/rules/falsifiable-checks.md`
|
|
132
|
+
- A zero-hit sweep is meaningless until the detector has found known instances on a ref where the defect still exists; a passing probe needs a deliberate bite control that MUST report a problem
|
|
133
|
+
- State each clean result's blind spot (presence vs. value, reachability, class completeness) — a negative result describes what the check can perceive, not the code
|
|
131
134
|
- Discover existing project scripts and tools before creating new ones
|
|
132
135
|
- Every verification must produce observable output -- a status code, a response body, a UI state, a test result
|
|
133
136
|
- Verification scripts must be runnable locally without CI/CD dependencies
|