@codyswann/lisa 2.213.0 → 2.215.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (196) hide show
  1. package/package.json +1 -1
  2. package/plugins/lisa/.claude-plugin/plugin.json +1 -1
  3. package/plugins/lisa/.codex-plugin/plugin.json +1 -1
  4. package/plugins/lisa/.codex-plugin/skills/lisa-codify-verification/SKILL.md +17 -0
  5. package/plugins/lisa/.codex-plugin/skills/lisa-drive-pr-to-merge/SKILL.md +54 -1
  6. package/plugins/lisa/.codex-plugin/skills/lisa-github-add-journey/SKILL.md +6 -6
  7. package/plugins/lisa/.codex-plugin/skills/lisa-github-journey/SKILL.md +13 -2
  8. package/plugins/lisa/.codex-plugin/skills/lisa-github-validate-issue/SKILL.md +9 -2
  9. package/plugins/lisa/.codex-plugin/skills/lisa-implement/SKILL.md +2 -2
  10. package/plugins/lisa/.codex-plugin/skills/lisa-jira-add-journey/SKILL.md +6 -6
  11. package/plugins/lisa/.codex-plugin/skills/lisa-jira-create/SKILL.md +5 -5
  12. package/plugins/lisa/.codex-plugin/skills/lisa-jira-journey/SKILL.md +15 -4
  13. package/plugins/lisa/.codex-plugin/skills/lisa-jira-validate-ticket/SKILL.md +9 -2
  14. package/plugins/lisa/.codex-plugin/skills/lisa-linear-add-journey/SKILL.md +7 -5
  15. package/plugins/lisa/.codex-plugin/skills/lisa-linear-create/SKILL.md +4 -4
  16. package/plugins/lisa/.codex-plugin/skills/lisa-linear-journey/SKILL.md +15 -4
  17. package/plugins/lisa/.codex-plugin/skills/lisa-linear-validate-issue/SKILL.md +9 -2
  18. package/plugins/lisa/.codex-plugin/skills/lisa-monitor/SKILL.md +1 -1
  19. package/plugins/lisa/.codex-plugin/skills/lisa-project-ideation/examples/idempotency-verification-harness.md +2 -2
  20. package/plugins/lisa/.codex-plugin/skills/lisa-tracker-add-journey/SKILL.md +1 -1
  21. package/plugins/lisa/.codex-plugin/skills/lisa-tracker-evidence/SKILL.md +1 -1
  22. package/plugins/lisa/.codex-plugin/skills/lisa-verification-lifecycle/SKILL.md +4 -4
  23. package/plugins/lisa/.codex-plugin/skills/lisa-verify/SKILL.md +1 -1
  24. package/plugins/lisa/rules/eager/observability-audit.md +1 -1
  25. package/plugins/lisa/rules/eager/verification.md +2 -1
  26. package/plugins/lisa/rules/reference/observability-audit.md +1 -1
  27. package/plugins/lisa/rules/reference/verification.md +43 -5
  28. package/plugins/lisa/skills/lisa-codify-verification/SKILL.md +17 -0
  29. package/plugins/lisa/skills/lisa-drive-pr-to-merge/SKILL.md +54 -1
  30. package/plugins/lisa/skills/lisa-github-add-journey/SKILL.md +6 -6
  31. package/plugins/lisa/skills/lisa-github-journey/SKILL.md +13 -2
  32. package/plugins/lisa/skills/lisa-github-validate-issue/SKILL.md +9 -2
  33. package/plugins/lisa/skills/lisa-implement/SKILL.md +2 -2
  34. package/plugins/lisa/skills/lisa-jira-add-journey/SKILL.md +6 -6
  35. package/plugins/lisa/skills/lisa-jira-create/SKILL.md +5 -5
  36. package/plugins/lisa/skills/lisa-jira-journey/SKILL.md +15 -4
  37. package/plugins/lisa/skills/lisa-jira-validate-ticket/SKILL.md +9 -2
  38. package/plugins/lisa/skills/lisa-linear-add-journey/SKILL.md +7 -5
  39. package/plugins/lisa/skills/lisa-linear-create/SKILL.md +4 -4
  40. package/plugins/lisa/skills/lisa-linear-journey/SKILL.md +16 -5
  41. package/plugins/lisa/skills/lisa-linear-validate-issue/SKILL.md +9 -2
  42. package/plugins/lisa/skills/lisa-monitor/SKILL.md +1 -1
  43. package/plugins/lisa/skills/lisa-project-ideation/examples/idempotency-verification-harness.md +2 -2
  44. package/plugins/lisa/skills/lisa-tracker-add-journey/SKILL.md +1 -1
  45. package/plugins/lisa/skills/lisa-tracker-evidence/SKILL.md +1 -1
  46. package/plugins/lisa/skills/lisa-verification-lifecycle/SKILL.md +4 -4
  47. package/plugins/lisa/skills/lisa-verify/SKILL.md +1 -1
  48. package/plugins/lisa-agy/plugin.json +1 -1
  49. package/plugins/lisa-agy/skills/lisa-codify-verification/SKILL.md +17 -0
  50. package/plugins/lisa-agy/skills/lisa-drive-pr-to-merge/SKILL.md +54 -1
  51. package/plugins/lisa-agy/skills/lisa-github-add-journey/SKILL.md +6 -6
  52. package/plugins/lisa-agy/skills/lisa-github-journey/SKILL.md +13 -2
  53. package/plugins/lisa-agy/skills/lisa-github-validate-issue/SKILL.md +9 -2
  54. package/plugins/lisa-agy/skills/lisa-implement/SKILL.md +2 -2
  55. package/plugins/lisa-agy/skills/lisa-jira-add-journey/SKILL.md +6 -6
  56. package/plugins/lisa-agy/skills/lisa-jira-create/SKILL.md +5 -5
  57. package/plugins/lisa-agy/skills/lisa-jira-journey/SKILL.md +15 -4
  58. package/plugins/lisa-agy/skills/lisa-jira-validate-ticket/SKILL.md +9 -2
  59. package/plugins/lisa-agy/skills/lisa-linear-add-journey/SKILL.md +7 -5
  60. package/plugins/lisa-agy/skills/lisa-linear-create/SKILL.md +4 -4
  61. package/plugins/lisa-agy/skills/lisa-linear-journey/SKILL.md +16 -5
  62. package/plugins/lisa-agy/skills/lisa-linear-validate-issue/SKILL.md +9 -2
  63. package/plugins/lisa-agy/skills/lisa-monitor/SKILL.md +1 -1
  64. package/plugins/lisa-agy/skills/lisa-project-ideation/examples/idempotency-verification-harness.md +2 -2
  65. package/plugins/lisa-agy/skills/lisa-tracker-add-journey/SKILL.md +1 -1
  66. package/plugins/lisa-agy/skills/lisa-tracker-evidence/SKILL.md +1 -1
  67. package/plugins/lisa-agy/skills/lisa-verification-lifecycle/SKILL.md +4 -4
  68. package/plugins/lisa-agy/skills/lisa-verify/SKILL.md +1 -1
  69. package/plugins/lisa-cdk/.claude-plugin/plugin.json +1 -1
  70. package/plugins/lisa-cdk/.codex-plugin/plugin.json +1 -1
  71. package/plugins/lisa-cdk-agy/plugin.json +1 -1
  72. package/plugins/lisa-cdk-copilot/.claude-plugin/plugin.json +1 -1
  73. package/plugins/lisa-cdk-cursor/.claude-plugin/plugin.json +1 -1
  74. package/plugins/lisa-copilot/.claude-plugin/plugin.json +1 -1
  75. package/plugins/lisa-copilot/rules/eager/observability-audit.md +1 -1
  76. package/plugins/lisa-copilot/rules/eager/verification.md +2 -1
  77. package/plugins/lisa-copilot/rules/reference/observability-audit.md +1 -1
  78. package/plugins/lisa-copilot/rules/reference/verification.md +43 -5
  79. package/plugins/lisa-copilot/skills/lisa-codify-verification/SKILL.md +17 -0
  80. package/plugins/lisa-copilot/skills/lisa-drive-pr-to-merge/SKILL.md +54 -1
  81. package/plugins/lisa-copilot/skills/lisa-github-add-journey/SKILL.md +6 -6
  82. package/plugins/lisa-copilot/skills/lisa-github-journey/SKILL.md +13 -2
  83. package/plugins/lisa-copilot/skills/lisa-github-validate-issue/SKILL.md +9 -2
  84. package/plugins/lisa-copilot/skills/lisa-implement/SKILL.md +2 -2
  85. package/plugins/lisa-copilot/skills/lisa-jira-add-journey/SKILL.md +6 -6
  86. package/plugins/lisa-copilot/skills/lisa-jira-create/SKILL.md +5 -5
  87. package/plugins/lisa-copilot/skills/lisa-jira-journey/SKILL.md +15 -4
  88. package/plugins/lisa-copilot/skills/lisa-jira-validate-ticket/SKILL.md +9 -2
  89. package/plugins/lisa-copilot/skills/lisa-linear-add-journey/SKILL.md +7 -5
  90. package/plugins/lisa-copilot/skills/lisa-linear-create/SKILL.md +4 -4
  91. package/plugins/lisa-copilot/skills/lisa-linear-journey/SKILL.md +16 -5
  92. package/plugins/lisa-copilot/skills/lisa-linear-validate-issue/SKILL.md +9 -2
  93. package/plugins/lisa-copilot/skills/lisa-monitor/SKILL.md +1 -1
  94. package/plugins/lisa-copilot/skills/lisa-project-ideation/examples/idempotency-verification-harness.md +2 -2
  95. package/plugins/lisa-copilot/skills/lisa-tracker-add-journey/SKILL.md +1 -1
  96. package/plugins/lisa-copilot/skills/lisa-tracker-evidence/SKILL.md +1 -1
  97. package/plugins/lisa-copilot/skills/lisa-verification-lifecycle/SKILL.md +4 -4
  98. package/plugins/lisa-copilot/skills/lisa-verify/SKILL.md +1 -1
  99. package/plugins/lisa-cursor/.claude-plugin/plugin.json +1 -1
  100. package/plugins/lisa-cursor/rules/observability-audit-reference.mdc +1 -1
  101. package/plugins/lisa-cursor/rules/observability-audit.mdc +1 -1
  102. package/plugins/lisa-cursor/rules/verification-reference.mdc +43 -5
  103. package/plugins/lisa-cursor/rules/verification.mdc +2 -1
  104. package/plugins/lisa-cursor/skills/lisa-codify-verification/SKILL.md +17 -0
  105. package/plugins/lisa-cursor/skills/lisa-drive-pr-to-merge/SKILL.md +54 -1
  106. package/plugins/lisa-cursor/skills/lisa-github-add-journey/SKILL.md +6 -6
  107. package/plugins/lisa-cursor/skills/lisa-github-journey/SKILL.md +13 -2
  108. package/plugins/lisa-cursor/skills/lisa-github-validate-issue/SKILL.md +9 -2
  109. package/plugins/lisa-cursor/skills/lisa-implement/SKILL.md +2 -2
  110. package/plugins/lisa-cursor/skills/lisa-jira-add-journey/SKILL.md +6 -6
  111. package/plugins/lisa-cursor/skills/lisa-jira-create/SKILL.md +5 -5
  112. package/plugins/lisa-cursor/skills/lisa-jira-journey/SKILL.md +15 -4
  113. package/plugins/lisa-cursor/skills/lisa-jira-validate-ticket/SKILL.md +9 -2
  114. package/plugins/lisa-cursor/skills/lisa-linear-add-journey/SKILL.md +7 -5
  115. package/plugins/lisa-cursor/skills/lisa-linear-create/SKILL.md +4 -4
  116. package/plugins/lisa-cursor/skills/lisa-linear-journey/SKILL.md +16 -5
  117. package/plugins/lisa-cursor/skills/lisa-linear-validate-issue/SKILL.md +9 -2
  118. package/plugins/lisa-cursor/skills/lisa-monitor/SKILL.md +1 -1
  119. package/plugins/lisa-cursor/skills/lisa-project-ideation/examples/idempotency-verification-harness.md +2 -2
  120. package/plugins/lisa-cursor/skills/lisa-tracker-add-journey/SKILL.md +1 -1
  121. package/plugins/lisa-cursor/skills/lisa-tracker-evidence/SKILL.md +1 -1
  122. package/plugins/lisa-cursor/skills/lisa-verification-lifecycle/SKILL.md +4 -4
  123. package/plugins/lisa-cursor/skills/lisa-verify/SKILL.md +1 -1
  124. package/plugins/lisa-expo/.claude-plugin/plugin.json +1 -1
  125. package/plugins/lisa-expo/.codex-plugin/plugin.json +1 -1
  126. package/plugins/lisa-expo-agy/plugin.json +1 -1
  127. package/plugins/lisa-expo-copilot/.claude-plugin/plugin.json +1 -1
  128. package/plugins/lisa-expo-cursor/.claude-plugin/plugin.json +1 -1
  129. package/plugins/lisa-harper-fabric/.claude-plugin/plugin.json +1 -1
  130. package/plugins/lisa-harper-fabric/.codex-plugin/plugin.json +1 -1
  131. package/plugins/lisa-harper-fabric-agy/plugin.json +1 -1
  132. package/plugins/lisa-harper-fabric-copilot/.claude-plugin/plugin.json +1 -1
  133. package/plugins/lisa-harper-fabric-cursor/.claude-plugin/plugin.json +1 -1
  134. package/plugins/lisa-nestjs/.claude-plugin/plugin.json +1 -1
  135. package/plugins/lisa-nestjs/.codex-plugin/plugin.json +1 -1
  136. package/plugins/lisa-nestjs-agy/plugin.json +1 -1
  137. package/plugins/lisa-nestjs-copilot/.claude-plugin/plugin.json +1 -1
  138. package/plugins/lisa-nestjs-cursor/.claude-plugin/plugin.json +1 -1
  139. package/plugins/lisa-openclaw/.claude-plugin/plugin.json +1 -1
  140. package/plugins/lisa-openclaw/.codex-plugin/plugin.json +1 -1
  141. package/plugins/lisa-openclaw-agy/plugin.json +1 -1
  142. package/plugins/lisa-openclaw-copilot/.claude-plugin/plugin.json +1 -1
  143. package/plugins/lisa-openclaw-cursor/.claude-plugin/plugin.json +1 -1
  144. package/plugins/lisa-phaser/.claude-plugin/plugin.json +1 -1
  145. package/plugins/lisa-phaser/.codex-plugin/plugin.json +1 -1
  146. package/plugins/lisa-phaser-agy/plugin.json +1 -1
  147. package/plugins/lisa-phaser-copilot/.claude-plugin/plugin.json +1 -1
  148. package/plugins/lisa-phaser-cursor/.claude-plugin/plugin.json +1 -1
  149. package/plugins/lisa-rails/.claude-plugin/plugin.json +1 -1
  150. package/plugins/lisa-rails/.codex-plugin/plugin.json +1 -1
  151. package/plugins/lisa-rails/.codex-plugin/skills/jira-journey/SKILL.md +1 -1
  152. package/plugins/lisa-rails/skills/jira-journey/SKILL.md +1 -1
  153. package/plugins/lisa-rails-agy/plugin.json +1 -1
  154. package/plugins/lisa-rails-agy/skills/jira-journey/SKILL.md +1 -1
  155. package/plugins/lisa-rails-copilot/.claude-plugin/plugin.json +1 -1
  156. package/plugins/lisa-rails-copilot/skills/jira-journey/SKILL.md +1 -1
  157. package/plugins/lisa-rails-cursor/.claude-plugin/plugin.json +1 -1
  158. package/plugins/lisa-rails-cursor/skills/jira-journey/SKILL.md +1 -1
  159. package/plugins/lisa-typescript/.claude-plugin/plugin.json +1 -1
  160. package/plugins/lisa-typescript/.codex-plugin/plugin.json +1 -1
  161. package/plugins/lisa-typescript-agy/plugin.json +1 -1
  162. package/plugins/lisa-typescript-copilot/.claude-plugin/plugin.json +1 -1
  163. package/plugins/lisa-typescript-cursor/.claude-plugin/plugin.json +1 -1
  164. package/plugins/lisa-wiki/.claude-plugin/plugin.json +1 -1
  165. package/plugins/lisa-wiki/.codex-plugin/plugin.json +1 -1
  166. package/plugins/lisa-wiki-agy/plugin.json +1 -1
  167. package/plugins/lisa-wiki-copilot/.claude-plugin/plugin.json +1 -1
  168. package/plugins/lisa-wiki-cursor/.claude-plugin/plugin.json +1 -1
  169. package/plugins/src/base/rules/eager/observability-audit.md +1 -1
  170. package/plugins/src/base/rules/eager/verification.md +2 -1
  171. package/plugins/src/base/rules/reference/observability-audit.md +1 -1
  172. package/plugins/src/base/rules/reference/verification.md +43 -5
  173. package/plugins/src/base/skills/lisa-codify-verification/SKILL.md +17 -0
  174. package/plugins/src/base/skills/lisa-drive-pr-to-merge/SKILL.md +54 -1
  175. package/plugins/src/base/skills/lisa-github-add-journey/SKILL.md +6 -6
  176. package/plugins/src/base/skills/lisa-github-journey/SKILL.md +13 -2
  177. package/plugins/src/base/skills/lisa-github-validate-issue/SKILL.md +9 -2
  178. package/plugins/src/base/skills/lisa-implement/SKILL.md +2 -2
  179. package/plugins/src/base/skills/lisa-jira-add-journey/SKILL.md +6 -6
  180. package/plugins/src/base/skills/lisa-jira-create/SKILL.md +5 -5
  181. package/plugins/src/base/skills/lisa-jira-journey/SKILL.md +15 -4
  182. package/plugins/src/base/skills/lisa-jira-validate-ticket/SKILL.md +9 -2
  183. package/plugins/src/base/skills/lisa-linear-add-journey/SKILL.md +7 -5
  184. package/plugins/src/base/skills/lisa-linear-create/SKILL.md +4 -4
  185. package/plugins/src/base/skills/lisa-linear-journey/SKILL.md +16 -5
  186. package/plugins/src/base/skills/lisa-linear-validate-issue/SKILL.md +9 -2
  187. package/plugins/src/base/skills/lisa-monitor/SKILL.md +1 -1
  188. package/plugins/src/base/skills/lisa-project-ideation/examples/idempotency-verification-harness.md +2 -2
  189. package/plugins/src/base/skills/lisa-tracker-add-journey/SKILL.md +1 -1
  190. package/plugins/src/base/skills/lisa-tracker-evidence/SKILL.md +1 -1
  191. package/plugins/src/base/skills/lisa-verification-lifecycle/SKILL.md +4 -4
  192. package/plugins/src/base/skills/lisa-verify/SKILL.md +1 -1
  193. package/plugins/src/rails/skills/jira-journey/SKILL.md +1 -1
  194. package/typescript/copy-overwrite/.github/GITHUB_ACTIONS.md +7 -3
  195. package/typescript/create-only/.github/workflows/claude-ci-auto-fix.yml +6 -2
  196. package/typescript/create-only/.github/workflows/claude-deploy-auto-fix.yml +4 -2
@@ -58,7 +58,7 @@ Items that change runtime behavior should include a `## Validation Journey` sect
58
58
 
59
59
  ### How to Write
60
60
 
61
- Design the journey based on the **change type**. Place `[EVIDENCE: name]` markers at key verification points.
61
+ Design the journey based on the **change type**. Place typed `[EVIDENCE: <artifact-type>: <name>]` markers at key verification points (types: `screenshot`, `recording`, `http-transcript`, `cli-output`, `log-snippet`, `db-query-output`, `perf-trace`, `test-run-log`, `deploy-log`, `state-dump` — see the `verification` rule).
62
62
 
63
63
  ```markdown
64
64
  ## Validation Journey
@@ -71,9 +71,9 @@ Design the journey based on the **change type**. Place `[EVIDENCE: name]` marker
71
71
  ### Steps
72
72
  1. Verify the current state before changes
73
73
  2. Apply the change (run migration, deploy, etc.)
74
- 3. Verify the expected new state [EVIDENCE: state-after-change]
75
- 4. Test error/edge cases [EVIDENCE: error-handling]
76
- 5. Verify rollback or cleanup if applicable [EVIDENCE: rollback-check]
74
+ 3. Verify the expected new state [EVIDENCE: http-transcript: state-after-change]
75
+ 4. Test error/edge cases [EVIDENCE: screenshot: error-state-rendered]
76
+ 5. Verify rollback or cleanup if applicable [EVIDENCE: db-query-output: rows-restored-after-rollback]
77
77
 
78
78
  ### Assertions
79
79
  - Describe what must be true after verification
@@ -1,12 +1,12 @@
1
1
  ---
2
2
  name: lisa-linear-journey
3
- description: "Parse a Linear Issue's Validation Journey section, execute the verification steps using appropriate tools (curl, test commands, database queries, Playwright), capture evidence at each [EVIDENCE: name] marker, and post to Linear + GitHub PR using the linear-evidence skill. Linear counterpart of lisa-jira-journey."
3
+ description: "Parse a Linear Issue's Validation Journey section, execute the verification steps using appropriate tools (curl, test commands, database queries, Playwright), capture evidence at each typed [EVIDENCE: <artifact-type>: <name>] marker, and post to Linear + GitHub PR using the linear-evidence skill. Linear counterpart of lisa-jira-journey."
4
4
  allowed-tools: ["Bash", "Read", "Glob", "Grep", "Skill"]
5
5
  ---
6
6
 
7
7
  # Linear Validation Journey
8
8
 
9
- Parse a Linear Issue's Validation Journey, execute the verification steps using the appropriate tools for the change type, capture evidence at each `[EVIDENCE: name]` marker, and post to Linear + GitHub PR.
9
+ Parse a Linear Issue's Validation Journey, execute the verification steps using the appropriate tools for the change type, capture evidence at each typed `[EVIDENCE: <artifact-type>: <name>]` marker, and post to Linear + GitHub PR.
10
10
 
11
11
  This skill is the destination of the `lisa-tracker-journey` shim when `tracker = "linear"`.
12
12
 
@@ -34,7 +34,7 @@ Reads `linear.workspace`, `linear.teamKey` from `.lisa.config.json` (with `.loca
34
34
  Fetch the Issue via `lisa-linear-access operation: get-issue` and extract the `## Validation Journey` section from the markdown description. Parse:
35
35
 
36
36
  - `### Prerequisites` — list of required services / env / setup
37
- - `### Steps` — numbered steps, each potentially containing `[EVIDENCE: name]` markers
37
+ - `### Steps` — numbered steps, each potentially containing typed `[EVIDENCE: <artifact-type>: <name>]` markers
38
38
  - `### Assertions` — what must be true after verification
39
39
 
40
40
  If the section is missing or has no steps, report `"No Validation Journey on <IDENTIFIER>. Run /linear-add-journey first."` and stop.
@@ -61,14 +61,25 @@ Execute each step sequentially. Determine the verification approach based on the
61
61
  - **Security fixes** → Reproduce exploit attempt, verify fix
62
62
  - **UI / frontend** → Playwright browser flow, capture screenshots / DOM state
63
63
 
64
- At each `[EVIDENCE: name]` marker, capture stdout / stderr to a numbered file:
64
+ At each typed `[EVIDENCE: <artifact-type>: <name>]` marker, capture an artifact **of the declared type** — the type is the contract, not a suggestion:
65
+
66
+ - `screenshot` / `recording` → an actual image/video file from the driven UI (Playwright, simulator), never a text description of what was seen
67
+ - `http-transcript` → the exact request (curl command or client call) plus the full response
68
+ - `cli-output` → the command plus stdout/stderr and exit code
69
+ - `log-snippet` → the correlated log lines pulled from the running system
70
+ - `db-query-output` → the query plus returned rows
71
+ - `perf-trace` → the benchmark/frame-timing/profiler output with methodology (device profile, dataset size)
72
+ - `test-run-log` → reporter output naming the spec and showing it ran and passed
73
+ - `deploy-log` / `state-dump` → the deployment/health-check output or observed-state JSON
74
+
75
+ A prose claim ("the error state rendered gracefully") satisfies no marker. Legacy untyped markers: infer the type from the step's action, capture accordingly, and note the inference. Write each artifact to a numbered file:
65
76
 
66
77
  #### Evidence Naming Convention
67
78
 
68
79
  `{NN}-{evidence-name}.{ext}`
69
80
 
70
81
  - `NN`: zero-padded sequential number (`01`, `02`, `03`...)
71
- - `evidence-name`: the value from `[EVIDENCE: name]`
82
+ - `evidence-name`: the `<name>` part of the typed marker
72
83
  - `ext`: `.txt` for plain output, `.json` for structured data
73
84
 
74
85
  Example:
@@ -200,9 +200,16 @@ An item with zero relations and no documented search: FAIL.
200
200
 
201
201
  #### S14 — Evidence manifest binding (leaf work units)
202
202
 
203
- When `issue_type ∈ {Bug, Task, Sub-task, Improvement}` AND `runtime_behavior_change = true`, the `## Validation Journey` must declare at least one `[EVIDENCE: name]` marker. Each marker name must be kebab-case and unique within the item. These markers are the work unit's **evidence manifest** — the exact, enumerated set of artifacts that must be captured and attached before the item may be closed (see the "Per-Work-Unit Evidence Contract" section of the `verification` rule, the Definition of Done in `verification-lifecycle`, and the evidence-manifest gate in `tracker-evidence`).
203
+ When `issue_type ∈ {Bug, Task, Sub-task, Improvement}` AND `runtime_behavior_change = true`, the `## Validation Journey` must declare at least one **typed** `[EVIDENCE: <artifact-type>: <name>]` marker. These markers are the work unit's **evidence manifest** — the exact, enumerated set of artifacts that must be captured and attached before the item may be closed (see the "Per-Work-Unit Evidence Contract" section of the `verification` rule, the Definition of Done in `verification-lifecycle`, and the evidence-manifest gate in `tracker-evidence`).
204
204
 
205
- FAIL when the Validation Journey is present but declares zero `[EVIDENCE: name]` markers, or when any marker name is empty, duplicated, or not kebab-case. A behavior-changing work unit SHOULD declare both a success marker and an error/edge marker; a journey with only one marker passes but the remediation should recommend adding the error/edge case.
205
+ Each marker must satisfy ALL of:
206
+
207
+ - `<artifact-type>` is one of the fixed taxonomy: `screenshot`, `recording`, `http-transcript`, `cli-output`, `log-snippet`, `db-query-output`, `perf-trace`, `test-run-log`, `deploy-log`, `state-dump`. (The legacy `[SCREENSHOT: name]` form is accepted as `screenshot`.)
208
+ - `<name>` is kebab-case and unique within the item.
209
+
210
+ **A marker names an artifact, not an assertion.** An untyped marker (`[EVIDENCE: load-failure-handled-gracefully]`) is an assertion label with nothing to capture and must FAIL, with a remediation that shows the typed transformation (e.g. → `[EVIDENCE: screenshot: load-failure-error-state]`, `[EVIDENCE: perf-trace: pipeline-load-tti]`).
211
+
212
+ FAIL when the Validation Journey is present but declares zero markers, when any marker is untyped or uses a type outside the taxonomy, or when any name is empty, duplicated, or not kebab-case. A behavior-changing work unit SHOULD declare both a success marker and an error/edge marker; a journey with only one marker passes but the remediation should recommend adding the error/edge case.
206
213
 
207
214
  This gate depends on S11. It is `N/A` for containers — a **Project** (the Epic equivalent), or any item with open child work (coordination containers, not work units) — and for leaf units with `runtime_behavior_change = false` (doc-only / config-only / type-only). If S11 fails because the Validation Journey is absent, S14 also FAILs (there is no manifest to bind) with remediation pointing back to `lisa-linear-add-journey`.
208
215
 
@@ -59,7 +59,7 @@ The **Rails** ops-specialist composes a different subset (`ops-run-local`, `ops-
59
59
  After report, file what was found — **only when run standalone**, never under `--report-only`/`--dry-run` and never when nested inside `lisa-verify` (which passes `--report-only`):
60
60
 
61
61
  - **Anomalies** (live signals over the conservative bar) → `Bug` leaves. **Gaps** (in-scope MISSING rubric dimensions) → `Task`/`Improvement` leaves.
62
- - Every ticket is filed via the vendor-neutral `lisa-tracker-write` shim with `build_ready: true` (never a vendor write skill directly), as a **single-repo leaf** stamped `repo:<current>`, with a real three-audience description, Gherkin AC, Target Backend Environment, and a Validation Journey + `EVIDENCE:` marker so it passes the `tracker-validate` gates.
62
+ - Every ticket is filed via the vendor-neutral `lisa-tracker-write` shim with `build_ready: true` (never a vendor write skill directly), as a **single-repo leaf** stamped `repo:<current>`, with a real three-audience description, Gherkin AC, Target Backend Environment, and a Validation Journey + typed `[EVIDENCE: <artifact-type>: <name>]` marker (e.g. `[EVIDENCE: log-snippet: alert-cleared]`) so it passes the `tracker-validate` gates.
63
63
  - **Idempotent:** embed the `<!-- lisa:monitor-finding: <fingerprint> -->` sentinel and search-before-create; never duplicate a live or just-resolved finding.
64
64
  - **Capped** at `max_candidates` (default 20), `core`/high-severity first; report how many were filed vs dropped.
65
65
  - **`--dry-run`** previews would-file tickets and creates nothing. **`--all-gaps`** widens gap filing to `recommended` tiers.
@@ -50,8 +50,8 @@ The harness performs the acceptance check in three phases:
50
50
 
51
51
  Capture the harness JSON output as:
52
52
 
53
- - `[EVIDENCE: marker-count-one]` from the first and second marker-count checks.
54
- - `[EVIDENCE: memory-recreated-after-rerun]` from the missing-memory variant.
53
+ - `[EVIDENCE: cli-output: marker-count-one]` from the first and second marker-count checks.
54
+ - `[EVIDENCE: cli-output: memory-recreated-after-rerun]` from the missing-memory variant.
55
55
 
56
56
  The run passes only when all reported counts are `1`, the issue URL is the same across phases, and
57
57
  `memoryRecreated` and `memoryFieldsRecorded` are `true` when `--memory-file` is supplied.
@@ -23,5 +23,5 @@ See the `config-resolution` rule for configuration and dispatch table.
23
23
 
24
24
  ## Rules
25
25
 
26
- - The Validation Journey content format is identical across all vendors (markdown sections with `[EVIDENCE: name]` markers). The only difference is how the section is appended — JIRA via `editJiraIssue` (Jira wiki markup), GitHub via `gh issue edit --body-file` (markdown), Linear via `save_issue` (markdown).
26
+ - The Validation Journey content format is identical across all vendors (markdown sections with typed `[EVIDENCE: <artifact-type>: <name>]` markers per the `verification` rule taxonomy). The only difference is how the section is appended — JIRA via `editJiraIssue` (Jira wiki markup), GitHub via `gh issue edit --body-file` (markdown), Linear via `save_issue` (markdown).
27
27
  - If the ticket already has a Validation Journey, the vendor skill reports it and stops. This shim does not retry.
@@ -27,7 +27,7 @@ See the `config-resolution` rule for configuration and dispatch table.
27
27
  - The GitHub `pr-assets` release lives on the implementation repo (the one with the PR), regardless of which tracker hosts the ticket/issue. All vendor skills upload there.
28
28
  - Never post evidence to a different ticket than the one named — `$ARGUMENTS` is the source of truth.
29
29
  - Never invent a verify-specific usage footer. Evidence artifact usage must flow through `lisa-usage-accounting`, preserve the canonical `## Lisa Usage` section, and surface `source: unavailable` explicitly when the runtime cannot provide trustworthy numbers.
30
- - **Evidence-manifest gate (leaf work units).** Before dispatching to a vendor skill that transitions the ticket, confirm `EVIDENCE_DIR` contains a non-empty artifact for every `[EVIDENCE: name]` marker declared in the ticket's Validation Journey. If any declared marker has no captured artifact (or only an empty one), stop and report the missing markers by name instead of posting — a leaf work unit (Bug / Task / Sub-task / Improvement) may not advance to its review/Done state with an unsatisfied manifest (see the "Per-Work-Unit Evidence Contract" in the `verification` rule). Epics / Stories / Spikes, and leaf units without a Validation Journey, are exempt.
30
+ - **Evidence-manifest gate (leaf work units).** Before dispatching to a vendor skill that transitions the ticket, confirm `EVIDENCE_DIR` contains a non-empty artifact **of the declared type** for every typed `[EVIDENCE: <artifact-type>: <name>]` marker declared in the ticket's Validation Journey — a `screenshot` marker needs an actual image, an `http-transcript` marker needs the request + response text, a `perf-trace` marker needs measured numbers; a prose claim satisfies nothing. If any declared marker has no captured artifact, an empty one, or one whose content/extension does not match its declared type, stop and report the offending markers by name instead of posting — a leaf work unit (Bug / Task / Sub-task / Improvement) may not advance to its review/Done state with an unsatisfied manifest (see the "Per-Work-Unit Evidence Contract" in the `verification` rule). Epics / Stories / Spikes, and leaf units without a Validation Journey, are exempt.
31
31
 
32
32
  ## UI Evidence Checklist (when work is UI-visible)
33
33
 
@@ -76,7 +76,7 @@ If auto-merge is enabled while the regression spec is still in flight, disable a
76
76
 
77
77
  After each empirical verification produces PASS evidence, invoke the `codify-verification` skill to encode the verification as an automated regression test. The manual proof becomes a repeatable check that catches future regressions.
78
78
 
79
- The `codify-verification` skill maps the verification type to the appropriate framework (Playwright for browser/UI, integration test for API/DB/auth, benchmark for performance, etc.), generates a deterministic test that asserts the same observable outcome the verification just confirmed, runs it in isolation to confirm PASS, and commits it in the same PR as the change.
79
+ The `codify-verification` skill maps the verification type to the appropriate framework (Playwright for browser/UI, integration test for API/DB/auth, benchmark for performance, etc.), generates a deterministic test that asserts the same observable outcome the verification just confirmed, runs it in isolation to confirm PASS, and commits it in the same PR as the change. For **frontend work**, codification is dual-runner: a Playwright spec in the project's Playwright test runner AND a Maestro flow in the Maestro test runner whenever the project supports Maestro (`.maestro/` directory, `maestro:test` script, or Maestro CI workflow) — both encoding the same verified journey, neither a substitute for the other.
80
80
 
81
81
  Codification is mandatory for every empirical verification type with one exception set: PR, Documentation, Deploy, and Investigate-Only spikes — those have inherently non-behavioral proof. For every other type, skipping codification is not allowed; if codification is genuinely impossible (e.g., the test framework does not exist and cannot be installed in scope), escalate via the Escalation Protocol rather than silently skipping.
82
82
 
@@ -236,7 +236,7 @@ Agents must follow this sequence unless explicitly instructed otherwise:
236
236
  8. Implement the change.
237
237
  9. Execute verification plan — run the actual system and observe results.
238
238
  10. Collect proof artifacts.
239
- 11. Codify — for each passing empirical verification, invoke `codify-verification` to encode it as a regression test (Playwright for UI, integration test for API/DB/auth, benchmark for performance, etc.) and commit the test in the same PR.
239
+ 11. Codify — for each passing empirical verification, invoke `codify-verification` to encode it as a regression test (Playwright for UI, integration test for API/DB/auth, benchmark for performance, etc.) and commit the test in the same PR. Frontend work codifies into every supported UI runner: Playwright spec + Maestro flow when the project supports Maestro (see the dual-runner section of `codify-verification`).
240
240
  12. Run spec conformance — build coverage matrix against the spec source (plan/ticket/issue), flag scope creep and untraceable changes, produce verdict.
241
241
  13. Summarize what changed, what was verified, what was codified, conformance verdict, and remaining risk.
242
242
  14. Label the result with a verification level.
@@ -250,7 +250,7 @@ Agents must follow this sequence unless explicitly instructed otherwise:
250
250
  3. **If verification fails**: Fix and re-run, don't mark complete
251
251
  4. **If verification blocked** (missing tools, services, etc.): Mark as blocked, not complete
252
252
  5. **Must not be dependent on CI/CD** if necessary, you may use local deploy methods found in the project manifest, but the verification methods must be listed in the pull request and therefore cannot be dependent on CI/CD completing
253
- 6. **Evidence manifest satisfied (leaf work units)**: For a leaf work unit (Bug / Task / Sub-task / Improvement) whose ticket carries a Validation Journey, do not mark the ticket complete or transition it out of in-progress until every `[EVIDENCE: name]` marker declared on the ticket has a corresponding captured, non-empty artifact attached to the ticket. A missing or empty artifact for any declared marker blocks completion exactly like a failed verification — fix and re-capture, or escalate; never close with an unsatisfied manifest. Epics / Stories / Spikes are exempt (coordination containers, not work units).
253
+ 6. **Evidence manifest satisfied (leaf work units)**: For a leaf work unit (Bug / Task / Sub-task / Improvement) whose ticket carries a Validation Journey, do not mark the ticket complete or transition it out of in-progress until every typed `[EVIDENCE: <artifact-type>: <name>]` marker declared on the ticket has a corresponding captured, non-empty artifact **of the declared type** attached to the ticket (an image for `screenshot`, request + response for `http-transcript`, measured output for `perf-trace`, …). A missing, empty, or wrong-type artifact for any declared marker blocks completion exactly like a failed verification — fix and re-capture, or escalate; never close with an unsatisfied manifest. Epics / Stories / Spikes are exempt (coordination containers, not work units).
254
254
  7. **No artifact-only completion for required runtime verification**: If empirical verification is required and cannot run because credentials are missing, do not mark the item done on artifact-only evidence. Exhaust the credential lookup order first; if still blocked, post the blocker comment, move the item to the configured blocked state, and apply the configured `needs-human` / `human-review` label.
255
255
 
256
256
  ---
@@ -357,7 +357,7 @@ A task is done only when:
357
357
  - Required verification surfaces and tooling surfaces are used or explicitly unavailable
358
358
  - Proof artifacts are captured
359
359
  - Every passing empirical verification is codified as a regression test (or has an explicit, documented skip reason from the allowed set)
360
- - For a leaf work unit, every `[EVIDENCE: name]` marker declared in its Validation Journey has a captured, non-empty artifact attached to the ticket (the evidence manifest is fully satisfied)
360
+ - For a leaf work unit, every typed `[EVIDENCE: <artifact-type>: <name>]` marker declared in its Validation Journey has a captured, non-empty artifact of the declared type attached to the ticket (the evidence manifest is fully satisfied)
361
361
  - Spec conformance verdict is `CONFORMS` (not `PARTIAL`, not `DIVERGES`)
362
362
  - Verification level is declared
363
363
  - Risks and gaps are documented
@@ -35,7 +35,7 @@ Treat the first successful lead-spawn request (or, on the Codex fallback, the fi
35
35
 
36
36
  Execute the **Verify** flow as defined in the `intent-routing` rule (loaded via the lisa plugin). The flow includes:
37
37
 
38
- 1. **Pre-flight: codification gate** — confirm that every passing local empirical verification on this branch was codified as a regression test (the Implement flow's codify step). If any verification has no committed test and no allowed skip reason (PR / Documentation / Deploy / Investigate-Only), invoke `codify-verification` now and amend the PR before shipping. A change cannot ship until its verifications are guarded.
38
+ 1. **Pre-flight: codification gate** — confirm that every passing local empirical verification on this branch was codified as a regression test (the Implement flow's codify step). If any verification has no committed test and no allowed skip reason (PR / Documentation / Deploy / Investigate-Only), invoke `codify-verification` now and amend the PR before shipping. For frontend work the gate is dual-runner: a Playwright spec AND, when the project supports Maestro (`.maestro/`, `maestro:test` script, or Maestro CI workflow), a Maestro flow for the same journey — a missing runner needs a recorded absence or a linked build-ready follow-up ticket, never a silent skip. A change cannot ship until its verifications are guarded.
39
39
  2. **Commit** any pending changes via `lisa-git-commit`
40
40
  3. **Push and PR** via `lisa-git-submit-pr`
41
41
  4. **Review loop** — handle CodeRabbit / human review comments via `lisa-pull-request-review`
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "lisa-cdk",
3
- "version": "2.213.0",
3
+ "version": "2.215.0",
4
4
  "description": "AWS CDK-specific plugin",
5
5
  "author": {
6
6
  "name": "Cody Swann"
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "lisa-cdk",
3
- "version": "2.213.0",
3
+ "version": "2.215.0",
4
4
  "description": "AWS CDK-specific Lisa plugin.",
5
5
  "author": {
6
6
  "name": "Cody Swann"
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "lisa-cdk",
3
- "version": "2.213.0",
3
+ "version": "2.215.0",
4
4
  "description": "AWS CDK-specific plugin",
5
5
  "author": {
6
6
  "name": "Cody Swann"
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "lisa-cdk",
3
- "version": "2.213.0",
3
+ "version": "2.215.0",
4
4
  "description": "AWS CDK-specific plugin",
5
5
  "author": {
6
6
  "name": "Cody Swann"
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "lisa-cdk",
3
- "version": "2.213.0",
3
+ "version": "2.215.0",
4
4
  "description": "AWS CDK-specific plugin",
5
5
  "author": {
6
6
  "name": "Cody Swann"
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "lisa",
3
- "version": "2.213.0",
3
+ "version": "2.215.0",
4
4
  "description": "Universal governance — agents, skills, commands, hooks, and rules for all projects",
5
5
  "author": {
6
6
  "name": "Cody Swann"
@@ -9,7 +9,7 @@ The **audit + file** arm of `lisa-monitor`. On top of its existing live-signal s
9
9
  - **Conservative by default.** Only high-signal anomalies (over the documented thresholds) and `core` missing dimensions are filed. `--all-gaps` widens gap filing to `recommended` tiers; nothing lowers the anomaly bar.
10
10
  - **Idempotent.** Every ticket carries a `<!-- lisa:monitor-finding: <fingerprint> -->` sentinel; search-before-create (including closed tickets) means a re-run never duplicates a live or just-resolved finding.
11
11
  - **Capped.** At most `monitor.maxCandidates` tickets per run (default 20), highest-severity first; report filed-vs-dropped (and list the dropped) — never silently truncate.
12
- - **Gate-passing.** Each ticket is a real authored artifact: three-audience description, Gherkin AC, single-repo, Target Backend Environment, and a Validation Journey with a unique kebab-case `[EVIDENCE: <name>]` marker — so `tracker-validate` (S1–S15) accepts it. A finding that cannot be made into a credible ticket is reported, not filed.
12
+ - **Gate-passing.** Each ticket is a real authored artifact: three-audience description, Gherkin AC, single-repo, Target Backend Environment, and a Validation Journey with a typed `[EVIDENCE: <artifact-type>: <name>]` marker (unique kebab-case name) — so `tracker-validate` (S1–S15) accepts it. A finding that cannot be made into a credible ticket is reported, not filed.
13
13
  - **Verify guard.** When `monitor` is the post-deploy step of `lisa-verify` it runs **report-only** — Verify invokes it as `lisa-monitor <env> --report-only`, so it never files there. Filing is a standalone-only action.
14
14
 
15
15
  Full reference (profile detection, the rubric table, anomaly thresholds, ticket templates, idempotency contract, the cap, dry-run/report-only semantics): [reference/observability-audit.md](../reference/observability-audit.md).
@@ -11,9 +11,10 @@
11
11
  - **Never claim success without runtime evidence.** "The code looks correct" is not evidence.
12
12
  - **If all you did was run tests, typecheck, and lint — you have NOT verified.**
13
13
  - **Before starting implementation, state your verification plan** — how you will USE the resulting software to prove it works. A plan that only lists `test`/`typecheck`/`lint` commands is not a plan. Do not begin until confirmed.
14
- - **After verifying empirically, codify it as a regression test** via the `codify-verification` skill — Playwright for UI, integration test for API/DB/auth, benchmark for performance. Codification is mandatory for every verification type except PR/Documentation/Deploy and Investigate-Only spikes.
14
+ - **After verifying empirically, codify it as a regression test** via the `codify-verification` skill — Playwright for UI, integration test for API/DB/auth, benchmark for performance. Codification is mandatory for every verification type except PR/Documentation/Deploy and Investigate-Only spikes. For **frontend work**, codification is dual-runner: a Playwright spec in the project's Playwright test runner AND a Maestro flow in the Maestro test runner whenever the project supports Maestro (`.maestro/` directory, `maestro:test` script, or Maestro CI workflow) — both encoding the same verified journey, neither a substitute for the other.
15
15
  - **The codified proof re-runs in CI.** Codified runtime verification lives where the project's e2e/Playwright tests live (`tests/e2e/**`) and runs as a required CI check for types with verification enforced — so the proof re-runs on every PR, not once by hand.
16
16
  - **Commit the evidence.** Write a durable artifact to `evidence/<ticket>/` — the acceptance criteria with per-criterion pass/fail + note, the observed state, screenshots/recording, and a `verdict.json`. The transient `.lisa/verification-status.json` is the session gate; `evidence/<ticket>/` is the committed proof.
17
+ - **Evidence markers are typed artifacts, not assertion labels.** A Validation Journey marker is `[EVIDENCE: <artifact-type>: <name>]` where the type (screenshot, recording, http-transcript, cli-output, log-snippet, db-query-output, perf-trace, test-run-log, deploy-log, state-dump) says HOW the proof is captured and the name says WHAT it proves. `[EVIDENCE: works-gracefully]` is a claim, not evidence — write `[EVIDENCE: screenshot: load-failure-error-state]`. Completion requires a captured artifact **of the declared type** per marker.
17
18
  - **Per-change verification is mandatory.** Every `feat`/`fix` adds or extends a verification (e2e) spec mapped to its acceptance criteria. The only exception is a genuinely non-behavioral change explicitly marked with the **logged** `verification-exempt` label — never a silent skip.
18
19
  - **Drive the playthrough with `/lisa:product-walkthrough`** (it walks the live product through a real browser); each project type plugs in its own drive mechanism (e.g. a Phaser game is driven through Playwright + an in-game verification test bridge that seeds RNG, reads state, injects input, and steps frames).
19
20
  - **Every PR must include reviewer replay steps** — the exact human steps to use the software and confirm the change works. Not test commands. If a reviewer can't reproduce from the PR description alone, the PR is incomplete.
@@ -73,7 +73,7 @@ Both finding types are filed through the vendor-neutral `lisa-tracker-write` shi
73
73
  - **Three-audience description** (S3): coding-assistant (technical: stack trace / Sentry link / occurrence count / the missing dimension and how to wire it), developer (where it surfaces, suspected cause/affected files, the fix skill to reach for), stakeholder (user/SLO impact).
74
74
  - **Gherkin acceptance criteria** (S4).
75
75
  - **Single repo** (S10): stamped `repo:<CURRENT_REPO>`; scope the description to this repo only.
76
- - **Target Backend Environment** (S8) + **Validation Journey** carrying at least one unique kebab-case `[EVIDENCE: <name>]` marker (S11/S14) — both anomaly fixes and gap wiring are runtime changes, so both need an env and a journey that proves the fix. The bracketed `[EVIDENCE: <name>]` form is what the validator scans for; a bare `EVIDENCE:` line fails S14. Prefer two markers (a success and an error/edge case).
76
+ - **Target Backend Environment** (S8) + **Validation Journey** carrying at least one typed `[EVIDENCE: <artifact-type>: <name>]` marker (S11/S14) — both anomaly fixes and gap wiring are runtime changes, so both need an env and a journey that proves the fix. The bracketed typed form is what the validator scans for; a bare `EVIDENCE:` line or an untyped assertion label (`[EVIDENCE: alert-fixed]`) fails S14 — for observability work the natural types are `log-snippet` and `state-dump`. Prefer two markers (a success and an error/edge case).
77
77
  - **Relationship search before write** (S13) — doubles as the dedup guard (next section).
78
78
  - **No parent/Epic required.** These are build-ready standalone leaves; S7's parent requirement is waived for a `build_ready: true` leaf (the leaf carve-out in `leaf-only-lifecycle` / `tracker-validate`). Do not fabricate an Epic parent.
79
79
  - A **priority** ordered by tier and severity: `core` gap / high-event anomaly → higher; `recommended` gap → lower.
@@ -100,12 +100,48 @@ Every change requires one or more verification types. Classify the change first,
100
100
 
101
101
  Every **leaf work unit** — an individually implementable ticket with no child tickets (issue types Bug, Task, Sub-task, Improvement) — that changes runtime behavior must declare, at creation time, the exact evidence that proves it is done. Epics, Stories, and Spikes are coordination containers, not work units: their evidence is the rollup of their children, so this contract does not apply to them.
102
102
 
103
- The declaration is not a separate field — it is the set of `[EVIDENCE: name]` markers in the work unit's **Validation Journey**. Those markers are the work unit's **evidence manifest**: an enumerated, named list of the artifacts a verifier must capture. The manifest binds both ends of the ticket lifecycle:
103
+ The declaration is not a separate field — it is the set of `[EVIDENCE: <artifact-type>: <name>]` markers in the work unit's **Validation Journey**. Those markers are the work unit's **evidence manifest**: an enumerated, typed list of the artifacts a verifier must capture. The manifest binds both ends of the ticket lifecycle:
104
104
 
105
- - **At creation** — the work unit cannot be written without a Validation Journey that names at least one `[EVIDENCE: name]` artifact (enforced by gate S14 in `tracker-validate` and the vendor `*-validate-*` skills). A behavior-changing unit should name both a success artifact and an error/edge artifact.
106
- - **At completion** — the work unit cannot be marked complete, nor transitioned to its review/Done state, until every `[EVIDENCE: name]` marker in its manifest has a captured, non-empty artifact attached to the ticket (enforced by the Task Completion Rules and Definition of Done in `verification-lifecycle`, and by the evidence-manifest gate in `tracker-evidence`). A manifest with a missing or empty artifact blocks completion exactly like a failed verification.
105
+ - **At creation** — the work unit cannot be written without a Validation Journey that declares at least one typed `[EVIDENCE: <artifact-type>: <name>]` artifact (enforced by gate S14 in `tracker-validate` and the vendor `*-validate-*` skills). A behavior-changing unit should declare both a success artifact and an error/edge artifact.
106
+ - **At completion** — the work unit cannot be marked complete, nor transitioned to its review/Done state, until every marker in its manifest has a captured, non-empty artifact **of the declared type** attached to the ticket (enforced by the Task Completion Rules and Definition of Done in `verification-lifecycle`, and by the evidence-manifest gate in `tracker-evidence`). A manifest with a missing, empty, or wrong-type artifact blocks completion exactly like a failed verification.
107
107
 
108
- The manifest is the single source of truth for "what evidence is required": authored once in the Validation Journey, enforced at write time, replayed during `tracker-journey`, and checked again before the ticket closes. There is no second list to keep in sync.
108
+ ### Marker grammar: the type is the evidence
109
+
110
+ **A marker names an artifact, not an assertion.** `[EVIDENCE: load-failure-handled-gracefully]` is a claim — there is nothing to capture, so it degenerates into a checkbox someone ticks. Evidence is empirical: a screenshot, a curl transcript, a log snippet, a frame-timing trace. The marker therefore carries two parts:
111
+
112
+ ```text
113
+ [EVIDENCE: <artifact-type>: <kebab-case-name>]
114
+ ```
115
+
116
+ - `<artifact-type>` — HOW the proof is captured, from the fixed taxonomy below.
117
+ - `<kebab-case-name>` — WHAT it proves, unique within the ticket.
118
+
119
+ Example transformation (the failure mode this grammar exists to prevent):
120
+
121
+ | Assertion label (invalid) | Typed artifact (valid) |
122
+ |---|---|
123
+ | `[EVIDENCE: pipeline-load-under-3s]` | `[EVIDENCE: perf-trace: pipeline-load-tti]` |
124
+ | `[EVIDENCE: long-column-virtualized]` | `[EVIDENCE: perf-trace: long-column-frame-timing]` |
125
+ | `[EVIDENCE: load-failure-handled-gracefully]` | `[EVIDENCE: screenshot: load-failure-error-state]` |
126
+
127
+ ### Artifact-type taxonomy (fixed set)
128
+
129
+ | Type | The captured artifact is |
130
+ |---|---|
131
+ | `screenshot` | an image of the observed UI/state (Playwright, simulator, device) |
132
+ | `recording` | a video of the playthrough |
133
+ | `http-transcript` | the exact request (curl command or client call) plus the full response — status, headers of interest, body |
134
+ | `cli-output` | the command plus its stdout/stderr and exit code |
135
+ | `log-snippet` | correlated log lines captured from the running system |
136
+ | `db-query-output` | the query plus the returned rows |
137
+ | `perf-trace` | benchmark / frame-timing / profiler output, with the methodology (device profile, dataset size) noted |
138
+ | `test-run-log` | reporter output naming the spec and showing it ran and passed |
139
+ | `deploy-log` | deployment output or a health-check response from the target environment |
140
+ | `state-dump` | machine-readable observed state (e.g. the `state.json` asserted against) |
141
+
142
+ Do not invent types inline; if none fits, propose extending this table. The legacy `[SCREENSHOT: name]` marker is equivalent to `[EVIDENCE: screenshot: name]`.
143
+
144
+ The manifest is the single source of truth for "what evidence is required": authored once in the Validation Journey, enforced at write time, replayed during `tracker-journey` (which captures each artifact **in its declared type**), and checked again before the ticket closes. There is no second list to keep in sync.
109
145
 
110
146
  ---
111
147
 
@@ -141,7 +177,7 @@ Verification **is** UAT — one gate, not two. This section makes the single
141
177
  verification process concrete so "an agent actually exercised the running
142
178
  software against the acceptance criteria" is durable and re-checkable, not a
143
179
  one-off claim. It builds on the **Per-Work-Unit Evidence Contract** above (the
144
- `[EVIDENCE: name]` manifest), adding three concrete requirements:
180
+ typed `[EVIDENCE: <artifact-type>: <name>]` manifest), adding three concrete requirements:
145
181
 
146
182
  **1. The codified proof re-runs in CI.** After local verification passes,
147
183
  `codify-verification` encodes it where the project's e2e/Playwright tests live
@@ -149,6 +185,8 @@ one-off claim. It builds on the **Per-Work-Unit Evidence Contract** above (the
149
185
  project type with verification **enforced**, it is a required check — so the proof
150
186
  re-runs on every PR instead of being proven once by hand.
151
187
 
188
+ For **frontend work**, codification is dual-runner: a Playwright spec in the project's Playwright test runner AND a Maestro flow in the Maestro test runner whenever the project supports Maestro (`.maestro/` directory, `maestro:test` script, or Maestro CI workflow) — both encoding the same verified journey, neither a substitute for the other. The dual-runner requirement is non-demotable: a missing runner is either a recorded absence (the project genuinely has no such harness) or a linked build-ready follow-up ticket — never a silent skip (see "Frontend dual-runner codification" in `codify-verification`).
189
+
152
190
  **2. Evidence is committed to the repo.** The named artifacts from the work
153
191
  unit's evidence manifest are committed under `evidence/<ticket>/` (in addition to
154
192
  any tracker attachment), so the proof lives with the code and is reviewable in the
@@ -57,6 +57,7 @@ Do NOT install a new framework if one already exists for the verification type.
57
57
  |---|---|
58
58
  | UI (web) | Playwright > Cypress > Selenium |
59
59
  | UI (mobile) | Maestro > Detox > Playwright (mobile emulation) |
60
+ | UI (frontend, project supports multiple runners) | **ALL supported UI runners** — see "Frontend dual-runner codification" below |
60
61
  | API | project's integration test runner (Vitest / Jest / RSpec / pytest) with HTTP client (supertest / fetch / faraday) |
61
62
  | Database | integration test with real DB + migrations applied |
62
63
  | Auth | API or UI test asserting role-gated access (multi-role coverage) |
@@ -71,6 +72,20 @@ Do NOT install a new framework if one already exists for the verification type.
71
72
 
72
73
  If the project lacks the preferred framework AND no acceptable substitute exists, escalate.
73
74
 
75
+ ### 2a. Frontend dual-runner codification (non-demotable)
76
+
77
+ For **frontend work** — any verification whose validation journey exercised a user-facing UI surface — codification is not one-runner-or-the-other. After the validation journey is complete and verified, the verified behavior MUST be codified in **every UI runner the project supports**:
78
+
79
+ 1. **A Playwright spec in the project's Playwright test runner** (where its web e2e tests live, e.g. `tests/e2e/**` / `e2e/**`) — required whenever the project has a Playwright (or equivalent web e2e) harness.
80
+ 2. **A Maestro flow in the project's Maestro test runner** — required whenever the project supports Maestro. Detect support by any of: a `.maestro/` directory (flows live in `.maestro/flows/`), a `maestro:test` script in `package.json`, or a Maestro CI workflow (e.g. `maestro-native-e2e`). Wire the new flow where the runner picks it up (`maestro test .maestro/flows`), tagging per the project's tier convention (e.g. `smoke`) when one exists.
81
+
82
+ Both artifacts encode the SAME verified journey — the Playwright spec drives the web surface, the Maestro flow drives the native surface. One is not a substitute for the other: they guard different platforms of the same behavior.
83
+
84
+ Permitted exits, mirroring the regression-spec rule in `lisa-implement` (never a silent skip, never "optional"):
85
+
86
+ - The project genuinely has no runner of that kind (no web e2e harness, or no Maestro support by the detection above) → record the checked locations and the absence in the codification evidence; that runner is N/A.
87
+ - A runner is supported but the flow/spec cannot be added or executed in this PR (genuine technical blocker) → create a linked build-ready follow-up ticket before merge, reference it from the PR and work item, and record the blocker — the same follow-up path as the regression-spec blocker.
88
+
74
89
  ### 3. Generate the test
75
90
 
76
91
  The generated test must:
@@ -103,6 +118,7 @@ requires a verification-spec delta on every behavioral change. See the
103
118
  Run only the new test, using whatever per-test invocation the project supports:
104
119
 
105
120
  - Playwright: `npx playwright test path/to/new.spec.ts`
121
+ - Maestro: `maestro test .maestro/flows/new-flow.yaml`
106
122
  - Vitest: `npx vitest run path/to/new.spec.ts`
107
123
  - Jest: `npx jest path/to/new.test.ts`
108
124
  - RSpec: `bundle exec rspec path/to/new_spec.rb`
@@ -139,6 +155,7 @@ Append to the verification report (or PR description):
139
155
  | # | Verification | Framework | Test file | Status |
140
156
  |---|--------------|-----------|-----------|--------|
141
157
  | 1 | <description> | Playwright | `e2e/checkout.spec.ts::displays order confirmation after checkout` | PASS |
158
+ | 2 | <same journey, native surface> | Maestro | `.maestro/flows/checkout-confirmation.yaml` | PASS |
142
159
  ```
143
160
 
144
161
  This evidence shows the verification is now guarded.
@@ -36,6 +36,43 @@ merges" loop. Other skills delegate here instead of re-implementing it. Runs
36
36
  Resolve `<owner>/<repo>` from `gh repo view --json nameWithOwner` (or the PR URL).
37
37
  Use plain `gh` + `git` so Claude and Codex execute identically.
38
38
 
39
+ ## 0. Take the babysitter lease
40
+
41
+ This skill is the branch's owner while it runs. Declare that ownership so the
42
+ CI auto-fix workflow (`reusable-claude-ci-auto-fix.yml`) stands down instead of
43
+ pushing competing fixes to the same branch (the single-writer rule):
44
+
45
+ ```bash
46
+ gh label create "lisa:babysitter-on-duty" \
47
+ --description "A drive-pr-to-merge session is actively driving this PR; CI auto-fix must stand down" \
48
+ --color FBCA04 || true # tolerate only already-exists; check the next step
49
+ gh pr edit <pr> --add-label "lisa:babysitter-on-duty"
50
+ gh pr view <pr> --json labels \
51
+ --jq '[.labels[].name] | contains(["lisa:babysitter-on-duty"])'
52
+ ```
53
+
54
+ Verify the final command prints `true` before driving. If the label could not
55
+ be attached (for example, no label-write permission), retry once; if it still
56
+ fails, surface a warning that the branch is unleased — the CI auto-fix
57
+ workflow may engage in parallel — and watch for its `claude-auto-fix-*` PR
58
+ per section 2f while driving.
59
+
60
+ The auto-fix workflow reads freshness from the label's most recent `labeled`
61
+ timeline event and treats stamps older than its TTL (default 90 minutes) as
62
+ stale. **Refresh the lease** whenever more than ~30 minutes have passed since
63
+ the last stamp while the watch loop is still running — a refresh is a
64
+ remove + re-add (re-adding an existing label does not create a new timeline
65
+ event):
66
+
67
+ ```bash
68
+ gh pr edit <pr> --remove-label "lisa:babysitter-on-duty"
69
+ gh pr edit <pr> --add-label "lisa:babysitter-on-duty"
70
+ ```
71
+
72
+ **Release the lease** (remove the label) at every terminal state — merged,
73
+ closed, or a hard block handed to a human. A crashed session that never
74
+ releases is why the TTL exists; do not rely on it as the normal release path.
75
+
39
76
  ## 1. Enable auto-merge
40
77
 
41
78
  Before enabling auto-merge, capture the live PR head and compare it to
@@ -74,7 +111,8 @@ gh pr view <pr> --json state,mergeStateStatus,mergeable,reviewDecision,statusChe
74
111
  ```
75
112
 
76
113
  Handle every blocker class; after any fix, re-poll and continue. Do not stop while
77
- the PR is still open and progress is possible.
114
+ the PR is still open and progress is possible. On each iteration, refresh the
115
+ babysitter lease if its last stamp is older than ~30 minutes (section 0).
78
116
 
79
117
  In **`on_blocker=report`** mode, only the mechanical step (a) and auto-merge enabling
80
118
  apply; for any of (b)–(e) do not act — classify the blocker and return per the input
@@ -150,6 +188,16 @@ Some org rulesets allow 0 approvals yet a bot `CHANGES_REQUESTED` still blocks
150
188
  auto-merge — dismissing the stale review after resolving all threads is what
151
189
  unblocks it.
152
190
 
191
+ ### f. Pending auto-fix PR into this branch
192
+ If an open PR from `claude-auto-fix-<headRefName>` targets this PR's head
193
+ branch (the CI auto-fix workflow engaged before this session took the lease),
194
+ adjudicate it: merge it into the head branch if the fix is correct and still
195
+ needed, otherwise close it and delete the side branch. Never leave it dangling
196
+ — it represents a competing writer's pending work. Merging it mutates the
197
+ driven branch, so treat it like any other push: disarm auto-merge first,
198
+ re-read `headRefOid`, reset `verify_commit` to the merged head, wait for that
199
+ head's checks to start, then re-enable auto-merge (section 1).
200
+
153
201
  ## 3. Merge and verify it actually shipped (ancestry check)
154
202
 
155
203
  Enabling auto-merge + green checks + resolved threads is **not** proof the merge
@@ -177,3 +225,8 @@ Loop until one of:
177
225
  needs design input, or genuine unresolved human objection (not a bot gate). Stop
178
226
  and report exactly what is blocking and what was already tried — never force the
179
227
  merge or weaken a gate to get past it.
228
+
229
+ At every terminal state, release the babysitter lease
230
+ (`gh pr edit <pr> --remove-label "lisa:babysitter-on-duty"`) so the CI
231
+ auto-fix workflow can take over as fixer of last resort if the branch goes
232
+ red later with nobody driving it.
@@ -58,7 +58,7 @@ Use Explore agents or read the codebase directly to understand which files are a
58
58
 
59
59
  ### Step 5: Draft the Validation Journey
60
60
 
61
- Compose the journey with `[EVIDENCE: name]` markers at key verification points:
61
+ Compose the journey with typed `[EVIDENCE: <artifact-type>: <name>]` markers at key verification points. The type says HOW the proof is captured (`screenshot`, `recording`, `http-transcript`, `cli-output`, `log-snippet`, `db-query-output`, `perf-trace`, `test-run-log`, `deploy-log`, `state-dump` — the fixed taxonomy in the `verification` rule); the name says WHAT it proves:
62
62
 
63
63
  ```markdown
64
64
  ## Validation Journey
@@ -69,9 +69,9 @@ Compose the journey with `[EVIDENCE: name]` markers at key verification points:
69
69
  ### Steps
70
70
  1. Verify current state before changes
71
71
  2. Apply the change
72
- 3. Verify expected new state [EVIDENCE: state-name]
73
- 4. Test error/edge cases [EVIDENCE: error-case]
74
- 5. Verify rollback if applicable [EVIDENCE: rollback]
72
+ 3. Verify expected new state [EVIDENCE: http-transcript: health-endpoint-200]
73
+ 4. Test error/edge cases [EVIDENCE: screenshot: invalid-input-error-state]
74
+ 5. Verify rollback if applicable [EVIDENCE: db-query-output: rows-restored-after-rollback]
75
75
 
76
76
  ### Assertions
77
77
  - Describe what must be true after verification
@@ -82,10 +82,10 @@ Compose the journey with `[EVIDENCE: name]` markers at key verification points:
82
82
  1. **2–5 evidence markers** — Focus on proving the change works and handles errors.
83
83
  2. **Concrete, runnable steps** — `Run \`curl -s localhost:3000/health | jq .status\`` not "Check the endpoint".
84
84
  3. **Include environment setup** — Database connection, running services, env vars.
85
- 4. **Evidence names in kebab-case** — `api-response`, `schema-check`, `rate-limit-hit`.
85
+ 4. **Markers are typed artifacts, not assertion labels** — `[EVIDENCE: <artifact-type>: <kebab-name>]`. `[EVIDENCE: load-failure-handled-gracefully]` names a claim with nothing to capture; write `[EVIDENCE: screenshot: load-failure-error-state]` or `[EVIDENCE: perf-trace: pipeline-load-tti]`. Names are kebab-case and unique within the ticket.
86
86
  5. **Assertions are measurable** — `Returns 200 with {status: ok}` not "API works correctly".
87
87
  6. **Cover happy path AND error path** — At minimum, one success and one failure marker.
88
- 7. **On a leaf work unit, the markers are binding** — For a Bug / Task / Sub-task / Improvement, every `[EVIDENCE: name]` here is the issue's evidence manifest: validation gate S14 requires at least one, and the issue cannot be closed until each named artifact is captured and attached (see the "Per-Work-Unit Evidence Contract" in the `verification` rule). Name only evidence you intend to capture — and name all of it.
88
+ 7. **On a leaf work unit, the markers are binding** — For a Bug / Task / Sub-task / Improvement, every typed `[EVIDENCE: <artifact-type>: <name>]` here is the issue's evidence manifest: validation gate S14 requires at least one, and the issue cannot be closed until each named artifact is captured **in its declared type** and attached (see the "Per-Work-Unit Evidence Contract" in the `verification` rule). Name only evidence you intend to capture — and name all of it.
89
89
 
90
90
  ### Step 6: Present to User for Approval
91
91
 
@@ -51,14 +51,25 @@ Execute each step sequentially. Determine the verification approach based on the
51
51
  - **Library/utility changes** → Run tests, capture output.
52
52
  - **Security fixes** → Reproduce exploit attempt, verify fix, capture output.
53
53
 
54
- At each `[EVIDENCE: name]` marker, capture stdout/stderr to a numbered file:
54
+ At each typed `[EVIDENCE: <artifact-type>: <name>]` marker, capture an artifact **of the declared type** — the type is the contract, not a suggestion:
55
+
56
+ - `screenshot` / `recording` → an actual image/video file from the driven UI (Playwright, simulator), never a text description of what was seen
57
+ - `http-transcript` → the exact request (curl command or client call) plus the full response
58
+ - `cli-output` → the command plus stdout/stderr and exit code
59
+ - `log-snippet` → the correlated log lines pulled from the running system
60
+ - `db-query-output` → the query plus returned rows
61
+ - `perf-trace` → the benchmark/frame-timing/profiler output with methodology (device profile, dataset size)
62
+ - `test-run-log` → reporter output naming the spec and showing it ran and passed
63
+ - `deploy-log` / `state-dump` → the deployment/health-check output or observed-state JSON
64
+
65
+ A prose claim ("the error state rendered gracefully") satisfies no marker. Legacy untyped markers: infer the type from the step's action, capture accordingly, and note the inference. Write each artifact to a numbered file:
55
66
 
56
67
  #### Evidence Naming Convention
57
68
 
58
69
  `{NN}-{evidence-name}.txt` (or `.json` for structured data):
59
70
 
60
71
  - `NN`: zero-padded sequential number (01, 02, 03...).
61
- - `evidence-name`: the value from `[EVIDENCE: <name>]` (kebab-case).
72
+ - `evidence-name`: the `<name>` part of the typed marker (kebab-case).
62
73
 
63
74
  Example:
64
75
 
@@ -196,9 +196,16 @@ An issue with zero links and no documented search: FAIL.
196
196
 
197
197
  #### S14 — Evidence manifest binding (leaf work units)
198
198
 
199
- When `issue_type ∈ {Bug, Task, Sub-task, Improvement}` AND `runtime_behavior_change = true`, the `## Validation Journey` must declare at least one `[EVIDENCE: name]` marker. Each marker name must be kebab-case and unique within the issue. These markers are the work unit's **evidence manifest** — the exact, enumerated set of artifacts that must be captured and attached before the issue may be closed (see the "Per-Work-Unit Evidence Contract" section of the `verification` rule, the Definition of Done in `verification-lifecycle`, and the evidence-manifest gate in `tracker-evidence`).
199
+ When `issue_type ∈ {Bug, Task, Sub-task, Improvement}` AND `runtime_behavior_change = true`, the `## Validation Journey` must declare at least one **typed** `[EVIDENCE: <artifact-type>: <name>]` marker. These markers are the work unit's **evidence manifest** — the exact, enumerated set of artifacts that must be captured and attached before the issue may be closed (see the "Per-Work-Unit Evidence Contract" section of the `verification` rule, the Definition of Done in `verification-lifecycle`, and the evidence-manifest gate in `tracker-evidence`).
200
200
 
201
- FAIL when the Validation Journey is present but declares zero `[EVIDENCE: name]` markers, or when any marker name is empty, duplicated, or not kebab-case. A behavior-changing work unit SHOULD declare both a success marker and an error/edge marker; a journey with only one marker passes but the remediation should recommend adding the error/edge case.
201
+ Each marker must satisfy ALL of:
202
+
203
+ - `<artifact-type>` is one of the fixed taxonomy: `screenshot`, `recording`, `http-transcript`, `cli-output`, `log-snippet`, `db-query-output`, `perf-trace`, `test-run-log`, `deploy-log`, `state-dump`. (The legacy `[SCREENSHOT: name]` form is accepted as `screenshot`.)
204
+ - `<name>` is kebab-case and unique within the issue.
205
+
206
+ **A marker names an artifact, not an assertion.** An untyped marker (`[EVIDENCE: load-failure-handled-gracefully]`) is an assertion label with nothing to capture and must FAIL, with a remediation that shows the typed transformation (e.g. → `[EVIDENCE: screenshot: load-failure-error-state]`, `[EVIDENCE: perf-trace: pipeline-load-tti]`).
207
+
208
+ FAIL when the Validation Journey is present but declares zero markers, when any marker is untyped or uses a type outside the taxonomy, or when any name is empty, duplicated, or not kebab-case. A behavior-changing work unit SHOULD declare both a success marker and an error/edge marker; a journey with only one marker passes but the remediation should recommend adding the error/edge case.
202
209
 
203
210
  This gate depends on S11. It is `N/A` for containers — an **Epic**, or any item with open child work (coordination containers, not work units) — and for leaf units with `runtime_behavior_change = false` (doc-only / config-only / type-only). If S11 fails because the Validation Journey is absent, S14 also FAILs (there is no manifest to bind) with remediation pointing back to `lisa-github-add-journey`.
204
211