@codyswann/lisa 2.263.0 → 2.265.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (107) hide show
  1. package/dist/core/upstream-evidence-manifest.d.ts.map +1 -1
  2. package/dist/core/upstream-evidence-manifest.js +13 -10
  3. package/dist/core/upstream-evidence-manifest.js.map +1 -1
  4. package/package.json +1 -1
  5. package/plugins/lisa/.claude-plugin/plugin.json +1 -1
  6. package/plugins/lisa/.codex-plugin/plugin.json +1 -1
  7. package/plugins/lisa/.codex-plugin/skills/lisa-github-evidence/SKILL.md +10 -0
  8. package/plugins/lisa/.codex-plugin/skills/lisa-jira-evidence/SKILL.md +10 -0
  9. package/plugins/lisa/.codex-plugin/skills/lisa-linear-evidence/SKILL.md +10 -0
  10. package/plugins/lisa/.codex-plugin/skills/lisa-security-review/SKILL.md +69 -2
  11. package/plugins/lisa/.codex-plugin/skills/lisa-security-zap-scan/SKILL.md +22 -2
  12. package/plugins/lisa/.codex-plugin/skills/lisa-tracker-evidence/SKILL.md +10 -3
  13. package/plugins/lisa/agents/security-specialist.md +19 -2
  14. package/plugins/lisa/rules/eager/claim-evidence-mapping.md +30 -3
  15. package/plugins/lisa/rules/reference/claim-evidence-mapping.md +155 -8
  16. package/plugins/lisa/rules/reference/verification.md +26 -0
  17. package/plugins/lisa/skills/lisa-github-evidence/SKILL.md +10 -0
  18. package/plugins/lisa/skills/lisa-jira-evidence/SKILL.md +10 -0
  19. package/plugins/lisa/skills/lisa-linear-evidence/SKILL.md +10 -0
  20. package/plugins/lisa/skills/lisa-security-review/SKILL.md +69 -2
  21. package/plugins/lisa/skills/lisa-security-zap-scan/SKILL.md +22 -2
  22. package/plugins/lisa/skills/lisa-tracker-evidence/SKILL.md +10 -3
  23. package/plugins/lisa-agy/agents/security-specialist.md +19 -2
  24. package/plugins/lisa-agy/plugin.json +1 -1
  25. package/plugins/lisa-agy/skills/lisa-github-evidence/SKILL.md +10 -0
  26. package/plugins/lisa-agy/skills/lisa-jira-evidence/SKILL.md +10 -0
  27. package/plugins/lisa-agy/skills/lisa-linear-evidence/SKILL.md +10 -0
  28. package/plugins/lisa-agy/skills/lisa-security-review/SKILL.md +69 -2
  29. package/plugins/lisa-agy/skills/lisa-security-zap-scan/SKILL.md +22 -2
  30. package/plugins/lisa-agy/skills/lisa-tracker-evidence/SKILL.md +10 -3
  31. package/plugins/lisa-cdk/.claude-plugin/plugin.json +1 -1
  32. package/plugins/lisa-cdk/.codex-plugin/plugin.json +1 -1
  33. package/plugins/lisa-cdk-agy/plugin.json +1 -1
  34. package/plugins/lisa-cdk-copilot/.claude-plugin/plugin.json +1 -1
  35. package/plugins/lisa-cdk-cursor/.claude-plugin/plugin.json +1 -1
  36. package/plugins/lisa-copilot/.claude-plugin/plugin.json +1 -1
  37. package/plugins/lisa-copilot/agents/security-specialist.agent.md +19 -2
  38. package/plugins/lisa-copilot/rules/eager/claim-evidence-mapping.md +30 -3
  39. package/plugins/lisa-copilot/rules/reference/claim-evidence-mapping.md +155 -8
  40. package/plugins/lisa-copilot/rules/reference/verification.md +26 -0
  41. package/plugins/lisa-copilot/skills/lisa-github-evidence/SKILL.md +10 -0
  42. package/plugins/lisa-copilot/skills/lisa-jira-evidence/SKILL.md +10 -0
  43. package/plugins/lisa-copilot/skills/lisa-linear-evidence/SKILL.md +10 -0
  44. package/plugins/lisa-copilot/skills/lisa-security-review/SKILL.md +69 -2
  45. package/plugins/lisa-copilot/skills/lisa-security-zap-scan/SKILL.md +22 -2
  46. package/plugins/lisa-copilot/skills/lisa-tracker-evidence/SKILL.md +10 -3
  47. package/plugins/lisa-cursor/.claude-plugin/plugin.json +1 -1
  48. package/plugins/lisa-cursor/agents/security-specialist.md +19 -2
  49. package/plugins/lisa-cursor/rules/claim-evidence-mapping-reference.mdc +155 -8
  50. package/plugins/lisa-cursor/rules/claim-evidence-mapping.mdc +30 -3
  51. package/plugins/lisa-cursor/rules/verification-reference.mdc +26 -0
  52. package/plugins/lisa-cursor/skills/lisa-github-evidence/SKILL.md +10 -0
  53. package/plugins/lisa-cursor/skills/lisa-jira-evidence/SKILL.md +10 -0
  54. package/plugins/lisa-cursor/skills/lisa-linear-evidence/SKILL.md +10 -0
  55. package/plugins/lisa-cursor/skills/lisa-security-review/SKILL.md +69 -2
  56. package/plugins/lisa-cursor/skills/lisa-security-zap-scan/SKILL.md +22 -2
  57. package/plugins/lisa-cursor/skills/lisa-tracker-evidence/SKILL.md +10 -3
  58. package/plugins/lisa-expo/.claude-plugin/plugin.json +1 -1
  59. package/plugins/lisa-expo/.codex-plugin/plugin.json +1 -1
  60. package/plugins/lisa-expo-agy/plugin.json +1 -1
  61. package/plugins/lisa-expo-copilot/.claude-plugin/plugin.json +1 -1
  62. package/plugins/lisa-expo-cursor/.claude-plugin/plugin.json +1 -1
  63. package/plugins/lisa-harper-fabric/.claude-plugin/plugin.json +1 -1
  64. package/plugins/lisa-harper-fabric/.codex-plugin/plugin.json +1 -1
  65. package/plugins/lisa-harper-fabric-agy/plugin.json +1 -1
  66. package/plugins/lisa-harper-fabric-copilot/.claude-plugin/plugin.json +1 -1
  67. package/plugins/lisa-harper-fabric-cursor/.claude-plugin/plugin.json +1 -1
  68. package/plugins/lisa-nestjs/.claude-plugin/plugin.json +1 -1
  69. package/plugins/lisa-nestjs/.codex-plugin/plugin.json +1 -1
  70. package/plugins/lisa-nestjs-agy/plugin.json +1 -1
  71. package/plugins/lisa-nestjs-copilot/.claude-plugin/plugin.json +1 -1
  72. package/plugins/lisa-nestjs-cursor/.claude-plugin/plugin.json +1 -1
  73. package/plugins/lisa-openclaw/.claude-plugin/plugin.json +1 -1
  74. package/plugins/lisa-openclaw/.codex-plugin/plugin.json +1 -1
  75. package/plugins/lisa-openclaw-agy/plugin.json +1 -1
  76. package/plugins/lisa-openclaw-copilot/.claude-plugin/plugin.json +1 -1
  77. package/plugins/lisa-openclaw-cursor/.claude-plugin/plugin.json +1 -1
  78. package/plugins/lisa-phaser/.claude-plugin/plugin.json +1 -1
  79. package/plugins/lisa-phaser/.codex-plugin/plugin.json +1 -1
  80. package/plugins/lisa-phaser-agy/plugin.json +1 -1
  81. package/plugins/lisa-phaser-copilot/.claude-plugin/plugin.json +1 -1
  82. package/plugins/lisa-phaser-cursor/.claude-plugin/plugin.json +1 -1
  83. package/plugins/lisa-rails/.claude-plugin/plugin.json +1 -1
  84. package/plugins/lisa-rails/.codex-plugin/plugin.json +1 -1
  85. package/plugins/lisa-rails-agy/plugin.json +1 -1
  86. package/plugins/lisa-rails-copilot/.claude-plugin/plugin.json +1 -1
  87. package/plugins/lisa-rails-cursor/.claude-plugin/plugin.json +1 -1
  88. package/plugins/lisa-typescript/.claude-plugin/plugin.json +1 -1
  89. package/plugins/lisa-typescript/.codex-plugin/plugin.json +1 -1
  90. package/plugins/lisa-typescript-agy/plugin.json +1 -1
  91. package/plugins/lisa-typescript-copilot/.claude-plugin/plugin.json +1 -1
  92. package/plugins/lisa-typescript-cursor/.claude-plugin/plugin.json +1 -1
  93. package/plugins/lisa-wiki/.claude-plugin/plugin.json +1 -1
  94. package/plugins/lisa-wiki/.codex-plugin/plugin.json +1 -1
  95. package/plugins/lisa-wiki-agy/plugin.json +1 -1
  96. package/plugins/lisa-wiki-copilot/.claude-plugin/plugin.json +1 -1
  97. package/plugins/lisa-wiki-cursor/.claude-plugin/plugin.json +1 -1
  98. package/plugins/src/base/agents/security-specialist.md +19 -2
  99. package/plugins/src/base/rules/eager/claim-evidence-mapping.md +30 -3
  100. package/plugins/src/base/rules/reference/claim-evidence-mapping.md +155 -8
  101. package/plugins/src/base/rules/reference/verification.md +26 -0
  102. package/plugins/src/base/skills/lisa-github-evidence/SKILL.md +10 -0
  103. package/plugins/src/base/skills/lisa-jira-evidence/SKILL.md +10 -0
  104. package/plugins/src/base/skills/lisa-linear-evidence/SKILL.md +10 -0
  105. package/plugins/src/base/skills/lisa-security-review/SKILL.md +69 -2
  106. package/plugins/src/base/skills/lisa-security-zap-scan/SKILL.md +22 -2
  107. package/plugins/src/base/skills/lisa-tracker-evidence/SKILL.md +10 -3
package/package.json CHANGED
@@ -106,7 +106,7 @@
106
106
  "form-data": ">=4.0.6"
107
107
  },
108
108
  "name": "@codyswann/lisa",
109
- "version": "2.263.0",
109
+ "version": "2.265.0",
110
110
  "description": "Claude Code governance framework that applies guardrails, guidance, and automated enforcement to projects",
111
111
  "main": "dist/index.js",
112
112
  "exports": {
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "lisa",
3
- "version": "2.263.0",
3
+ "version": "2.265.0",
4
4
  "description": "Universal governance — agents, skills, commands, hooks, and rules for all projects",
5
5
  "author": {
6
6
  "name": "Cody Swann"
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "lisa",
3
- "version": "2.263.0",
3
+ "version": "2.265.0",
4
4
  "description": "Universal governance: agents, skills, commands, hooks, and rules for all projects.",
5
5
  "author": {
6
6
  "name": "Cody Swann"
@@ -24,6 +24,16 @@ Upload captured evidence and generated templates to the GitHub PR description an
24
24
  - `comment.md` — GitHub markdown body for both the issue comment and the PR description's `## Evidence` section.
25
25
  - (Optional) `comment.txt` — kept for parity with the JIRA path; not used here.
26
26
 
27
+ ## Comment-body preflight (required)
28
+
29
+ Before posting or updating anything, check the evidence body (`comment.md`, and `comment.txt` where this skill uses it):
30
+
31
+ - It contains a `## Not established` heading. That heading is **never omitted and never blank** — when nothing is outstanding it still renders `None outstanding — reviewed`; otherwise it names, in plain operator language, what the verification did not prove.
32
+ - The accompanying verdict carries `not_established_reviewed: true` (the list may be empty; the flag may never be omitted).
33
+ - It contains a `## Artifact identity` heading carrying **values, not placeholders** — the repository, the `head_sha` the verification observed, the `environment`, and per artifact its `sha256` digest and `captured_at`. **Refuse to post** a body whose identity heading is absent or unpopulated, or whose recorded `artifact_head_sha` disagrees with the verdict's `artifact.head_sha` — report the evidence id and **both SHAs**. Definition: the `claim-evidence-mapping` rule.
34
+
35
+ If either is missing, **refuse to post**: stop and report the missing Not-established review to the caller instead of publishing. Composing the body is `lisa-tracker-evidence`'s job (see its UI Evidence Checklist); this skill only refuses to publish one that omits the section. The section is defined by the `claim-evidence-mapping` rule and generalizes `lisa-improve-harness`'s required, never-empty `Known limits` field.
36
+
27
37
  ## Workflow
28
38
 
29
39
  1. **Resolve refs**
@@ -42,6 +42,16 @@ Upload captured evidence and generated templates to GitHub PR description and JI
42
42
  - `comment.txt` — JIRA wiki markup (generated by `generate-templates.py`)
43
43
  - `comment.md` — GitHub markdown (generated by `generate-templates.py`)
44
44
 
45
+ ## Comment-body preflight (required)
46
+
47
+ Before posting or updating anything, check the evidence body (`comment.md`, and `comment.txt` where this skill uses it):
48
+
49
+ - It contains a `## Not established` heading. That heading is **never omitted and never blank** — when nothing is outstanding it still renders `None outstanding — reviewed`; otherwise it names, in plain operator language, what the verification did not prove.
50
+ - The accompanying verdict carries `not_established_reviewed: true` (the list may be empty; the flag may never be omitted).
51
+ - It contains a `## Artifact identity` heading carrying **values, not placeholders** — the repository, the `head_sha` the verification observed, the `environment`, and per artifact its `sha256` digest and `captured_at`. **Refuse to post** a body whose identity heading is absent or unpopulated, or whose recorded `artifact_head_sha` disagrees with the verdict's `artifact.head_sha` — report the evidence id and **both SHAs**. Definition: the `claim-evidence-mapping` rule.
52
+
53
+ If either is missing, **refuse to post**: stop and report the missing Not-established review to the caller instead of publishing. Composing the body is `lisa-tracker-evidence`'s job (see its UI Evidence Checklist); this skill only refuses to publish one that omits the section. The section is defined by the `claim-evidence-mapping` rule and generalizes `lisa-improve-harness`'s required, never-empty `Known limits` field.
54
+
45
55
  ## Usage
46
56
 
47
57
  ```bash
@@ -41,6 +41,16 @@ The caller must produce:
41
41
 
42
42
  If any of these are missing, stop and report.
43
43
 
44
+ ## Comment-body preflight (required)
45
+
46
+ Before posting or updating anything, check the evidence body (`comment.md`, and `comment.txt` where this skill uses it):
47
+
48
+ - It contains a `## Not established` heading. That heading is **never omitted and never blank** — when nothing is outstanding it still renders `None outstanding — reviewed`; otherwise it names, in plain operator language, what the verification did not prove.
49
+ - The accompanying verdict carries `not_established_reviewed: true` (the list may be empty; the flag may never be omitted).
50
+ - It contains a `## Artifact identity` heading carrying **values, not placeholders** — the repository, the `head_sha` the verification observed, the `environment`, and per artifact its `sha256` digest and `captured_at`. **Refuse to post** a body whose identity heading is absent or unpopulated, or whose recorded `artifact_head_sha` disagrees with the verdict's `artifact.head_sha` — report the evidence id and **both SHAs**. Definition: the `claim-evidence-mapping` rule.
51
+
52
+ If either is missing, **refuse to post**: stop and report the missing Not-established review to the caller instead of publishing. Composing the body is `lisa-tracker-evidence`'s job (see its UI Evidence Checklist); this skill only refuses to publish one that omits the section. The section is defined by the `claim-evidence-mapping` rule and generalizes `lisa-improve-harness`'s required, never-empty `Known limits` field.
53
+
44
54
  ## Phase 1 — Resolve Linear Issue
45
55
 
46
56
  1. Parse the identifier from `$ARGUMENTS`.
@@ -16,6 +16,57 @@ Identify vulnerabilities, evaluate threats, and recommend mitigations for code c
16
16
  5. **Check auth/authz** -- are access controls properly enforced for new endpoints or features?
17
17
  6. **Review dependencies** -- do new dependencies introduce known vulnerabilities?
18
18
 
19
+ ## The impact-or-exploitability bar
20
+
21
+ Severity is **earned, not pattern-matched**. Every security-shaped finding is classified
22
+ mechanically, before it is written up:
23
+
24
+ | Field | What it holds |
25
+ |-------|---------------|
26
+ | `reproducer` | an evidence ref of a kind that reaches the claim's boundary, or `none` |
27
+ | `impact` | a bounded impact/exploitability statement (who can do what, to what data, under what preconditions), or `unproven` |
28
+ | `reason` | one line saying why the finding landed in its bucket |
29
+
30
+ **The bar:** a finding is **proven** only when it carries **both** a reproducer **and** a bounded
31
+ impact statement. **Missing either ⇒ unproven.** No other input changes the bucket.
32
+
33
+ What counts as a reaching reproducer is defined by the `claim-evidence-mapping` contract (**BCE-1**,
34
+ #1835), not here: an injection claim at the `http-api` boundary needs an `http-transcript`; a UI
35
+ claim needs a `screenshot` or `recording`. A passing unit `test-run-log` reaches `code-unit` only and
36
+ never discharges either.
37
+
38
+ **Each field stands on its own.** The two halves are recorded independently: a finding with a bounded
39
+ impact but no reproducer keeps its impact statement verbatim and only `reproducer` reads `none`; a
40
+ finding with a reproducer but no bounded impact keeps the evidence ref and only `impact` reads
41
+ `unproven`. **Never overwrite** a field you actually have with a missing-value placeholder — the
42
+ `reason` line names which half is missing, and the surviving half is the head start the next reviewer
43
+ needs.
44
+
45
+ ## The two buckets — conservative by default
46
+
47
+ Findings render in two clearly-labeled buckets: **Security (proven)** and **Security (unproven)**.
48
+
49
+ A reproducer-less finding **stays in the security section**, labeled `unproven` with its reason. It
50
+ is **never auto-demoted** to a `maintenance` bucket and it is **not removed** from the report —
51
+ under-reporting a real vulnerability is the worse failure, so the conservative default keeps it
52
+ visible where a security reader looks.
53
+
54
+ **Single policy point.** The unproven bucket's label is the only thing an owner may change:
55
+ `security.review.unprovenBucket` in `.lisa.config.json`, default `security-unproven`. An owner who
56
+ prefers true demotion sets it to a maintenance label; the finding then renders under that bucket and
57
+ **no other classification logic changes** — the bar, the fields, and the reasons are identical.
58
+
59
+ Write both buckets in operator voice (`factory-model` rule 5): a person who does not code reads this
60
+ at the gate. "Anyone who can reach the search box can read other customers' orders — reproduced with
61
+ the request transcript below" is usable; "possible SQLi in handler" is not.
62
+
63
+ This bar governs *code-review* security findings. Dependency CVE remediation keeps its own decision
64
+ ladder in the `security-audit-handling` rule — cite it, do not restate or fork it.
65
+
66
+ Bucketing is **advisory** — it shapes the report, it does not block a merge — on the same terms as
67
+ the boundary checks, which stay reporting-only until `verification.gate.enforceBoundaries` is `true`
68
+ in `.lisa.config.json`.
69
+
19
70
  ## Output Format
20
71
 
21
72
  Structure findings as:
@@ -41,17 +92,33 @@ Structure findings as:
41
92
  - [ ] No XSS vectors in user-facing output
42
93
  - [ ] Dependencies free of known CVEs
43
94
 
44
- ### Vulnerabilities Found
45
- - [vulnerability] -- where in the code, how to prevent
95
+ ### Security (proven)
96
+ - [finding] -- where in the code, how to prevent
97
+ - reproducer: [evidence ref, e.g. evidence/<ticket>/http-transcript-01.txt]
98
+ - impact: [who can do what, to what data, under what preconditions]
99
+ - reason: reproducer + bounded impact
100
+
101
+ ### Security (unproven)
102
+ - [finding] -- where in the code, how to prevent
103
+ - reproducer: [evidence ref if one exists, else `none`]
104
+ - impact: [bounded statement if one exists, else `unproven`]
105
+ - reason: [which half is missing -- e.g. "impact bounded, but never reproduced"]
106
+ -- kept in the security section, not demoted
46
107
 
47
108
  ### Recommendations
48
109
  - [recommendation] -- priority (critical/warning/suggestion)
49
110
  ```
50
111
 
112
+ Rename the unproven heading only when `security.review.unprovenBucket` is set to something other
113
+ than `security-unproven`; everything else stays as written.
114
+
51
115
  ## Rules
52
116
 
53
117
  - Focus on the specific changes proposed, not a full security audit of the entire codebase
54
118
  - Flag only real risks -- do not invent hypothetical threats for internal tooling with no user input
119
+ - Classify every finding against the bar before writing it up; never leave a finding unbucketed
120
+ - Never silently drop or downgrade a finding out of the security section -- `unproven` is the
121
+ conservative landing spot, and the reason line says why
55
122
  - Prioritize OWASP Top 10 vulnerabilities
56
123
  - If the changes are purely internal (config, refactoring, docs), report "No security concerns" and explain why
57
124
  - Always check `.gitleaksignore` patterns to understand what secrets scanning is already in place
@@ -22,10 +22,30 @@ Run a ZAP baseline security scan against the local application.
22
22
  - After the scan completes, read `zap-report.html` (or `zap-report.md` for text)
23
23
  - Summarize findings:
24
24
  - Total number of alerts by risk level (High, Medium, Low, Informational)
25
- - List each Medium+ finding with its rule ID, name, and recommended fix
25
+ - **Every alert reaches classification** -- High, Medium, Low, and Informational alike. Risk
26
+ level orders the summary; it never filters it. **Nothing is dropped before classification**,
27
+ so no alert can leave the report unclassified. Medium+ alerts are listed first, in full (rule
28
+ ID, name, recommended fix); Low/Informational alerts are still listed, bucketed, and given a
29
+ `reason`, even when compressed to one line each.
26
30
  - Categorize findings as "infrastructure-level" (fix at CDN/proxy) vs "application-level" (fix in code)
27
31
 
28
- 4. **Handle failures**:
32
+ 4. **Apply the impact-or-exploitability bar** -- the same bar the `lisa-security-review` skill
33
+ defines; follow that skill, do not restate it. A ZAP alert is not a reproducer by itself: the
34
+ alert names a pattern, not an exercised impact path.
35
+ - **Security (proven)** -- the alert carries a reproducer **and** a bounded impact statement. The
36
+ reproducer counts only if its evidence kind **reaches the claim's boundary** under the
37
+ `claim-evidence-mapping` contract (BCE-1, #1835): a ZAP request/response transcript is an
38
+ `http-transcript` and reaches the `http-api` boundary only. An alert whose claim is about
39
+ rendered UI (`browser`) or persisted state (`data`) needs evidence at *that* boundary -- a
40
+ transcript never proves it.
41
+ - **Security (unproven)** -- everything else, each with a one-line `reason` (typically
42
+ "alert only, no reproducer / no bounded impact", or "transcript does not reach the claim's
43
+ boundary"). Unproven alerts are **not dropped** and not demoted out of the security summary --
44
+ they render in the unproven bucket so a reader still sees them.
45
+ - Rename the unproven heading only if `security.review.unprovenBucket` is set to something other
46
+ than `security-unproven`; no other classification changes.
47
+
48
+ 5. **Handle failures**:
29
49
  - If the scan failed, explain what failed and suggest concrete remediation steps
30
50
 
31
51
  ## Execution
@@ -28,6 +28,10 @@ See the `config-resolution` rule for configuration and dispatch table.
28
28
  - Never post evidence to a different ticket than the one named — `$ARGUMENTS` is the source of truth.
29
29
  - Never invent a verify-specific usage footer. Evidence artifact usage must flow through `lisa-usage-accounting`, preserve the canonical `## Lisa Usage` section, and surface `source: unavailable` explicitly when the runtime cannot provide trustworthy numbers.
30
30
  - **Evidence-manifest gate (leaf work units).** Before dispatching to a vendor skill that transitions the ticket, confirm `EVIDENCE_DIR` contains a non-empty artifact **of the declared type** for every typed `[EVIDENCE: <artifact-type>: <name>]` marker declared in the ticket's Validation Journey — a `screenshot` marker needs an actual image, an `http-transcript` marker needs the request + response text, a `perf-trace` marker needs measured numbers; a prose claim satisfies nothing. If any declared marker has no captured artifact, an empty one, or one whose content/extension does not match its declared type, stop and report the offending markers by name instead of posting — a leaf work unit (Bug / Task / Sub-task / Improvement) may not advance to its review/Done state with an unsatisfied manifest (see the "Per-Work-Unit Evidence Contract" in the `verification` rule). Epics / Stories / Spikes, and leaf units without a Validation Journey, are exempt.
31
+ - **Claim↔boundary binding (S14 upgrade).** Satisfying the manifest by *type* is not enough: each `[EVIDENCE: <artifact-type>: <name>]` marker is also cited for a **claim**, and the marker's artifact type must **reach the boundary that claim declares** per the `claim-evidence-mapping` rule's taxonomy. A `browser` claim (user-visible UI behavior) is reached by `screenshot` / `recording` and never by a unit `test-run-log`; an `http-api` claim needs an `http-transcript`; a `deploy-health` claim needs a `deploy-log` and no pre-deploy artifact. On a mismatch, report the offending marker **by name** together with the claim, its boundary, and the **required evidence kinds** for that boundary — e.g. `[EVIDENCE: test-run-log: unit-suite] cited for claim AC-2 [boundary browser] — required evidence kinds: screenshot, recording`. Swapping in a marker of a reaching kind with a real captured artifact satisfies the gate. **Advisory-first:** until `verification.gate.enforceBoundaries` is `true` in `.lisa.config.json` (the same ratchet flag the Stop-hook gate reads, default `false`), a boundary mismatch is reported to the operator but does not block the post; once ratcheted on it refuses the post exactly like a missing or wrong-type artifact. Missing/empty/wrong-type artifacts keep refusing the post regardless of the flag.
32
+ - **The "Not established" section is required.** Before posting, confirm `evidence/comment.md` (and `comment.txt` where the vendor uses it) contains a `## Not established` heading and that the verdict carries `not_established_reviewed: true`. The heading is **never omitted and never blank**: with nothing outstanding it still renders `None outstanding — reviewed`; otherwise it lists, in plain operator language, each thing the verification did not prove — boundaries not exercised, environments not tested, behavior consciously out of scope. A comment with no such heading, a heading with nothing under it, or a verdict whose `not_established_reviewed` flag is absent is **refused**: stop and report the missing Not-established review instead of posting. (The list may be empty; the flag may never be omitted.) Definition and exemplars live in the `claim-evidence-mapping` rule; this generalizes `lisa-improve-harness`'s required, never-empty `Known limits` field.
33
+ - **Artifact identity is required, with values (S14 extension).** Before posting, confirm the comment body carries a `## Artifact identity` heading populated with **values, not placeholders**: the `repository`, the `head_sha` the verification observed, the `environment`, and for each committed artifact its `sha256` digest and `captured_at`. Then check both identity failures. **`artifact_mismatch`** — an evidence entry whose recorded `artifact_head_sha` differs from the verdict's `artifact.head_sha`: refuse the post and report the offending evidence id together with **both SHAs** (the one the artifact was captured at and the one the verdict claims), e.g. `EV-2 captured at 4f1c9ab but artifact.head_sha is 9de0c31`. **`evidence_digest_mismatch`** — recompute the `sha256` of each committed evidence file on read; if the bytes no longer match the recorded digest (or the file is absent), stop and report the evidence id by name. Never summarize either as "verification failed" — name the field, the ids, and both SHAs so a non-engineer can read it at the gate. **Advisory-first:** until `verification.gate.enforceBoundaries` is `true` in `.lisa.config.json`, an identity failure is reported to the operator but does not block the post; once ratcheted on it refuses the post exactly like a missing artifact. Definition: the `claim-evidence-mapping` rule.
34
+ - **Merge-race reconciliation (at completion, not at post).** The evidence's pinned `head_sha` is reconciled against the merged head using the ancestry + deploy-run definition of "what shipped" that `lisa-drive-pr-to-merge` already owns — cite that skill, never re-implement or restate its checks here. Pre-merge evidence counts for the merge commit only when its head is a parent of the merge, ancestry alone never justifies reporting shipped, and on a mismatch verification re-runs against the merged head before completion is declared.
31
35
  - **Evidence references are not manifest entries.** Extract obligations using the exact `[EVIDENCE:` prefix (and the legacy local `[SCREENSHOT:` form). Exclude both the canonical `[EVIDENCE-REF: <work-item-ref> | <artifact-type>: <kebab-case-name>]` and the Lisa 2.223.0 legacy alias `[EVIDENCE-REF: <tracker-ref>: <artifact-type>: <kebab-case-name>]` from artifact lookup, missing-artifact reporting, and duplicate-name checks: either belongs to another work item and cannot satisfy this item's S14 gate. A runtime-changing leaf with references but no local claiming marker must be rejected before dispatch.
32
36
 
33
37
  ## UI Evidence Checklist (when work is UI-visible)
@@ -47,8 +51,11 @@ The checklist is tracker-agnostic — the same shape works on JIRA, GitHub Issue
47
51
  5. **"What this shows" section.** Tailor to ticket type:
48
52
  - **Bug repro:** state plainly whether the bug reproduces or not, and the most likely 1–2 reasons their retest still failed (different env, native app vs. web, stuck backend row, etc.).
49
53
  - **Feature/UX completion:** state plainly which acceptance criteria each screenshot covers, and call out any deferred or out-of-scope surface explicitly so QA/PM doesn't have to infer.
50
- 6. **"What would help me confirm" (bug) / "How to QA" (feature) section.** Concrete actionable retest steps with the exact selection criteria (e.g., "pick a record whose Status column shows `—`, not `Processing` or `Pending Review`").
51
- 7. **Explicit invitation to be corrected.** End with a line like *"If any of the steps I listed are different from what you expected / actually did, please tell me explicitly which step I got wrong."* Non-optional small differences (which record, which device, exact tap order, expected behavior) change everything, and naming the door open short-circuits ticket bounce-loops.
52
- 8. **Workflow transition** is the vendor skill's job, not yoursit'll move the ticket per the configured tracker (JIRA: Reassign to reporter for bug repro / move to the configured review status when one exists, otherwise leave it in `claimed`; GitHub: direct `claimed` configured `done` after a successful build; Linear: equivalent state). You don't transition manually.
54
+ 6. **"Artifact identity" and "Not established" sections (both required).** Right after "What this shows":
55
+ - `## Artifact identity` what the evidence was collected against, as **values, not placeholders**: the repository, the `head_sha` the verification observed, the `environment`, and per artifact its `sha256` digest and `captured_at`. A heading rendered with `<unknown>` in place of a SHA is not identity.
56
+ - `## Not established` — **never omitted, never blank.** List in plain language what this verification did *not* prove: boundaries not exercised (e.g. "the persisted order row was never queried"), environments not tested (e.g. "checked on desktop Chrome only not mobile Safari"), and behavior consciously out of scope (e.g. "refunds were not touched or tested"). Name the specific thing, not a category. With nothing outstanding, the heading still renders a single line: `None outstanding reviewed`. Never state a quality check as proof of behavior "unit tests pass" belongs here as a limit, not above as evidence.
57
+ 7. **"What would help me confirm" (bug) / "How to QA" (feature) section.** Concrete actionable retest steps with the exact selection criteria (e.g., "pick a record whose Status column shows `—`, not `Processing` or `Pending Review`").
58
+ 8. **Explicit invitation to be corrected.** End with a line like *"If any of the steps I listed are different from what you expected / actually did, please tell me explicitly which step I got wrong."* Non-optional — small differences (which record, which device, exact tap order, expected behavior) change everything, and naming the door open short-circuits ticket bounce-loops.
59
+ 9. **Workflow transition** is the vendor skill's job, not yours — it'll move the ticket per the configured tracker (JIRA: Reassign to reporter for bug repro / move to the configured review status when one exists, otherwise leave it in `claimed`; GitHub: direct `claimed` → configured `done` after a successful build; Linear: equivalent state). You don't transition manually.
53
60
 
54
61
  **Why this format:** It (a) gives the reporter a frame-by-frame they can compare against, (b) avoids the JIRA image-collapse failure mode while still working everywhere else, (c) names the most plausible discrepancies up front so the loop short-circuits, (d) explicitly opens the door to being corrected so tickets don't bounce on assumed alignment. The same mechanics that resolve a stuck bug ticket also give QA an unambiguous handoff for a freshly-built feature.
@@ -35,13 +35,30 @@ Structure your findings as:
35
35
  - [ ] No XSS vectors in user-facing output
36
36
  - [ ] Dependencies free of known CVEs
37
37
 
38
- ### Vulnerabilities Found
39
- - [vulnerability] -- where in the code, how to prevent
38
+ ### Security (proven)
39
+ - [finding] -- where in the code, how to prevent
40
+ - reproducer: [evidence ref]
41
+ - impact: [who can do what, to what data, under what preconditions]
42
+ - reason: reproducer + bounded impact
43
+
44
+ ### Security (unproven)
45
+ - [finding] -- where in the code, how to prevent
46
+ - reproducer: [evidence ref if one exists, else `none`]
47
+ - impact: [bounded statement if one exists, else `unproven`]
48
+ - reason: [which half is missing -- e.g. "impact bounded, but never reproduced"]
49
+ -- kept in the security section, not demoted
40
50
 
41
51
  ### Recommendations
42
52
  - [recommendation] -- priority (critical/warning/suggestion)
43
53
  ```
44
54
 
55
+ A finding is **proven** only with both a reproducer evidence ref and a bounded impact statement;
56
+ missing either, it stays **unproven** inside the security section. Record the two halves
57
+ independently -- keep whichever one you have and let the `reason` name the missing half; never
58
+ overwrite a real value with a placeholder. The full bar, the per-finding fields, and the
59
+ `security.review.unprovenBucket` policy point live in the `security-review` skill -- follow it, do
60
+ not restate it.
61
+
45
62
  ## Rules
46
63
 
47
64
  - Focus on the specific changes proposed, not a full security audit of the entire codebase
@@ -37,9 +37,36 @@ rule 5).
37
37
  A claim carries three fields — `claim_id`, `boundary`, and `required_evidence_kinds` — named here so
38
38
  every downstream surface uses one spelling. This ticket only writes the contract down; the schema and
39
39
  gate that make these fields executable ship with **BCE-2 (#1836)** — do not assume that surface is
40
- present in this branch. A claim with no reaching evidence is **Not established** (defined fully in
41
- **BCE-3 (#1837)**), an artifact's identity is pinned in **BCE-4 (#1838)**, and the conservative
42
- security-bucket default is set in **BCE-5 (#1839)** — each named here, defined there.
40
+ present in this branch.
41
+
42
+ ## Security buckets (conservative by default)
43
+
44
+ A security finding is **proven** only with both a reproducer of a reaching kind and a bounded
45
+ impact/exploitability statement; missing either, it renders **unproven** with its reason and **stays
46
+ in the security section** — never auto-demoted to maintenance. The label is one policy point,
47
+ `security.review.unprovenBucket` (default `security-unproven`); the procedure lives in the
48
+ `lisa-security-review` skill.
49
+
50
+ ## Artifact identity (pinned, never assumed)
51
+
52
+ A claim applies only to the artifact its evidence was collected against. Every verdict pins
53
+ `artifact.head_sha` — the commit the run observed — and every evidence entry pins the
54
+ `artifact_head_sha` in force when it was captured, its `sha256` content digest, and `captured_at`.
55
+ Evidence collected on a pre-merge head is valid for the merge commit **only** when that head is a
56
+ parent of the merge, per the ancestry + deploy-run definition of "what shipped" that
57
+ `lisa-drive-pr-to-merge` already owns — cite it, never write a second one. A mismatched SHA or a
58
+ recomputed digest that disagrees fails loudly, naming both SHAs / the evidence id, and verification
59
+ re-runs against the merged head before completion is declared. Full definition: the reference body.
60
+
61
+ ## Not established (required, never omitted)
62
+
63
+ A claim with no reaching evidence is **Not established**, and every report says so out loud. Each
64
+ evidence comment and verdict carries a `Not established` section listing what was *not* proved —
65
+ boundaries not exercised, environments not tested, behavior consciously out of scope. It is never
66
+ omitted and never blank: with nothing outstanding it still renders `None outstanding — reviewed`, and
67
+ `not_established_reviewed` attests the list was reviewed even when the list itself is empty. This
68
+ generalizes the required, never-empty `Known limits` field of `lisa-improve-harness`. Full definition
69
+ and operator-voice exemplars: the reference body.
43
70
 
44
71
  ## No behavior change; degrade, never block
45
72
 
@@ -80,12 +80,11 @@ defines the names; it stores nothing:
80
80
  | `required_evidence_kinds` | the evidence kind(s) that reach that boundary, from the `verification` artifact-type set |
81
81
 
82
82
  A claim whose `required_evidence_kinds` has no captured, reaching artifact is **Not established** —
83
- the concept is named here and defined fully, with its evidence templates, in **BCE-3 (#1837)**; do
84
- not assume that section is present in this branch. Artifact identity what makes two captured
85
- artifacts the same or different is pinned in **BCE-4 (#1838)**, and the conservative default
86
- bucket for a security-sensitive claim is set in **BCE-5 (#1839)**. Each is named here as the field
87
- BCE-2's schema will carry; none is defined by this contract. Each ships with that ticket — do not
88
- assume its section is present in this branch.
83
+ defined in full, with its evidence templates, in the section below. Artifact identity — what makes
84
+ two captured artifacts the same or different — is defined in the *Artifact identity* section below
85
+ (shipped by BCE-4, #1838). The bucket a security-sensitive claim lands in is defined in the *Security
86
+ buckets* section below (shipped by BCE-5, #1839). Where such a sibling surface is not installed in a
87
+ given branch, name what you can and continue.
89
88
 
90
89
  ### Worked example
91
90
 
@@ -114,13 +113,161 @@ Claim: "The service is deployed and healthy."
114
113
  response from the target environment.
115
114
  ```
116
115
 
116
+ ## Artifact identity — what the evidence was collected against
117
+
118
+ A claim reaches only as far as its evidence's *kind*; it applies only to the *artifact* that evidence
119
+ was collected against. **Artifact identity is what makes two captured artifacts the same or
120
+ different**: one repository, at one commit, in one environment, at one moment. Evidence that does not
121
+ say which artifact it observed is not evidence about anything in particular — it silently transfers
122
+ to whatever ships next, which is exactly how an auto-merge race ships code no verification ever
123
+ touched.
124
+
125
+ Identity is carried in two places, and they must agree:
126
+
127
+ | Where | Field | What it pins |
128
+ |---|---|---|
129
+ | `artifact` (once per verdict) | `repository` | the `owner/repo` the run observed |
130
+ | | `base_sha` | the base the change was measured against |
131
+ | | `head_sha` | **the commit the verification actually observed** — required for a v2 pass |
132
+ | | `build_id` | the build/run the evidence came from, where one exists |
133
+ | | `environment` | where it ran (local, preview, staging, production) |
134
+ | | `observed_at` | when the run observed it (ISO-8601 UTC) |
135
+ | `evidence[]` (per artifact) | `artifact_head_sha` | the `head_sha` in force **when that artifact was captured** |
136
+ | | `sha256` | content digest of the committed evidence file |
137
+ | | `captured_at` | when that artifact was captured (ISO-8601 UTC) |
138
+
139
+ `head_sha` pins the *build*; `sha256` pins the *bytes*. Together they answer both identity questions:
140
+ "which artifact was this collected against" and "is this still the artifact that was collected". The
141
+ `sha256` + commit-ref discipline is the same one Lisa's upstream-evidence manifest already uses — it
142
+ is cited as prior art here, not reinvented.
143
+
144
+ ### Two identity failures, both loud
145
+
146
+ - **`artifact_mismatch`** — an `evidence[]` entry whose `artifact_head_sha` differs from the verdict's
147
+ `artifact.head_sha`. The evidence describes a different build than the one the verdict claims. The
148
+ identity check fails **loudly, naming both SHAs** (the evidence's and the verdict's) so an operator
149
+ can see which build each half is talking about.
150
+ - **`evidence_digest_mismatch`** — on read, recompute the `sha256` of each committed evidence file. If
151
+ the bytes no longer match the recorded digest, the artifact has changed since it was recorded: the
152
+ check fails, it names the evidence id, and blocks completion. A digest that cannot be recomputed
153
+ (file absent) is the same failure.
154
+
155
+ Neither failure is ever summarized as "verification failed". Name the field, the ids, and both SHAs —
156
+ a person who does not code reads this at the gate (`factory-model` rule 5). Both checks are
157
+ **advisory-first**: reported to the operator but non-blocking until
158
+ `verification.gate.enforceBoundaries` is `true` in `.lisa.config.json` — the same ratchet flag the
159
+ boundary checks and the Stop-hook gate ride.
160
+
161
+ ### The merge race — one definition, two guards
162
+
163
+ "The artifact that shipped" is already defined once, in `lisa-drive-pr-to-merge`: the **ancestry**
164
+ check (is the verified commit an ancestor of the merged base branch, and is the merge commit's parent
165
+ that verified head rather than a stale one) **plus** the **deploy-run** check (a deploy/release run
166
+ actually fired for the merge SHA or an including descendant). Cite that definition; never write a
167
+ second one. Identity reconciliation is the same definition read from the evidence side:
168
+
169
+ - Evidence collected on a **pre-merge head** is valid for the merge commit **only when that head is a
170
+ parent of the merge** — i.e. it satisfies that skill's ancestry check against the merged head. If a
171
+ late commit raced past the merge, the merged head is not the verified one and the evidence does not
172
+ transfer.
173
+ - Ancestry alone is never enough. That skill's rule stands unchanged here: **never report shipped on
174
+ ancestry alone** — the deploy-run check must also pass before completion is declared.
175
+ - On mismatch (`artifact.head_sha` ≠ the reconciled merge SHA), completion is not declared. Flag the
176
+ mismatch naming both SHAs and **re-run verification against the merged head**; the re-run's verdict
177
+ is what may declare completion.
178
+
179
+ Two guards, one definition — so they cannot disagree about what shipped.
180
+
181
+ ## Security buckets — the impact-or-exploitability bar
182
+
183
+ A security-shaped claim is a claim like any other: it reaches only as far as its evidence. Applied to
184
+ security findings (shipped by BCE-5, #1839), that yields two buckets and one bar:
185
+
186
+ - **Security (proven)** — the finding carries **both** a `reproducer` (an evidence ref of a kind that
187
+ reaches the claim's boundary, per the taxonomy above — e.g. an `http-transcript` for an injection
188
+ claim at `http-api`) **and** a bounded `impact`/exploitability statement.
189
+ - **Security (unproven)** — **missing either**. It keeps a one-line `reason` and **stays in the
190
+ security section**: a reproducer-less finding is **never auto-demoted** to a `maintenance` bucket,
191
+ because under-reporting a real vulnerability is the worse failure.
192
+
193
+ The unproven bucket's label is the single configurable policy point —
194
+ `security.review.unprovenBucket` in `.lisa.config.json`, default `security-unproven`. An owner who
195
+ prefers true demotion flips that one value; no other classification logic changes. Like the boundary
196
+ checks, bucketing is **advisory** — it shapes the report, not the merge — until
197
+ `verification.gate.enforceBoundaries` is `true`.
198
+
199
+ The full procedure — per-finding fields, report shape, ZAP alignment — lives in the
200
+ `lisa-security-review` skill (with `lisa-security-zap-scan` citing it); dependency-CVE remediation
201
+ keeps its separate ladder in `security-audit-handling`. Cite those slugs; do not restate them here.
202
+
203
+ ## The "Not established" section — required, never omitted
204
+
205
+ Every report that asserts something was verified also states, in the same breath, **what it did not
206
+ establish**. This is one section, it is **required**, and it is **never omitted and never blank** —
207
+ on the evidence comment, in the committed `evidence/<ticket>/verdict.json`, and in the
208
+ `verification-status.json` verdict BCE-2's gate reads.
209
+
210
+ **Where it appears and what it says**
211
+
212
+ | Surface | Shape |
213
+ |---|---|
214
+ | Evidence comment (tracker + PR `## Evidence` section) | a `## Not established` heading, one plain-language bullet per item |
215
+ | `evidence/<ticket>/verdict.json` | `not_established: []` plus `not_established_reviewed: true` |
216
+ | `.lisa/verification-status.json` (schema v2) | per-claim `not_established[]` plus the top-level `not_established_reviewed` flag |
217
+
218
+ **The empty case is not the omitted case.** A verification that genuinely left nothing unproved still
219
+ renders the heading, with the single line:
220
+
221
+ ```text
222
+ ## Not established
223
+
224
+ None outstanding — reviewed
225
+ ```
226
+
227
+ An absent heading, or a heading with nothing under it, is a defect — it is indistinguishable from
228
+ never having asked the question. The list may be empty; the section may not be blank.
229
+
230
+ **Machine-readable semantics (as shipped by BCE-2, #1836).** `not_established` is the list — it may
231
+ be empty. `not_established_reviewed` is the boolean attestation that the list was actually
232
+ reviewed — **the flag may never be omitted**. That asymmetry is the whole mechanism: an empty list
233
+ plus a present flag means "we looked and found nothing outstanding"; an absent flag means nobody
234
+ looked. The Stop-hook gate treats an absent flag as a v2 contract violation, reported to stderr and
235
+ **advisory** until `verification.gate.enforceBoundaries` is `true` in `.lisa.config.json` (the same
236
+ ratchet flag the boundary checks ride); evidence surfaces refuse the post on the same terms.
237
+
238
+ **What belongs under the heading** — written in operator voice (`factory-model` rule 5: a person who
239
+ does not code reads this at the gate), not in engineering shorthand:
240
+
241
+ - **Boundaries not exercised** — a claim's boundary that no captured artifact reached. *"The checkout
242
+ button was proved in the browser; the order's persisted row was never queried, so the `data`
243
+ boundary is not established."*
244
+ - **Environments not tested** — where it was and was not run. *"Checked on production Chrome at
245
+ 1440×900 only. Not checked on mobile Safari, and not checked against the staging database."*
246
+ - **Claims consciously out of scope** — deliberately excluded behavior, named so nobody infers it.
247
+ *"Refunds were not touched or tested; this change covers new orders only."*
248
+ - **Anything a green quality check might be mistaken for proving.** *"Unit tests pass for the submit
249
+ handler. That establishes the code-unit boundary only — it is not evidence the button works."*
250
+
251
+ Each item names the thing, not a category: "not tested on mobile Safari" is usable at a gate; "some
252
+ environments untested" is not.
253
+
254
+ ### Philosophical precedent for this section
255
+
256
+ This generalizes `lisa-improve-harness`'s **`Known limits`** field — a required, never-empty line on
257
+ every result record. That skill says it plainly:
258
+ *"A record with nothing in it is invalid on its face"* — because a single-trajectory loop always has
259
+ limits. The same is true of any verification:
260
+ it ran somewhere, on something, once. `Known limits` (one record) and `not_established` (every claim
261
+ in the factory) are the same discipline with the same never-empty rule; read either and you should
262
+ recognize the other.
263
+
117
264
  ## Philosophical precedent
118
265
 
119
266
  This generalizes the **bounded-claim discipline** of `lisa-improve-harness`: one trajectory supports
120
267
  one trajectory's claim, and a result record may claim only what its cited evidence reaches. Here the
121
268
  same discipline is applied to every claim in the factory — a claim reaches exactly as far as the
122
- *kind* of evidence behind it, and no further. BCE-3 generalizes the *Not established* half of that
123
- discipline into a first-class report state.
269
+ *kind* of evidence behind it, and no further. The *Not established* half of that discipline is a
270
+ first-class report state — see the required, never-omitted section above.
124
271
 
125
272
  ## No behavior change; degrade, never block
126
273
 
@@ -151,6 +151,32 @@ Do not invent types inline; if none fits, propose extending this table. The lega
151
151
 
152
152
  The manifest is the single source of truth for "what evidence is required": authored once in the Validation Journey, enforced at write time, replayed during `tracker-journey` (which captures each artifact **in its declared type**), and checked again before the ticket closes. There is no second list to keep in sync.
153
153
 
154
+ ### Every evidence surface names what it did NOT establish
155
+
156
+ An evidence comment that lists only what passed is unreadable at a gate: a journey that skipped an edge state looks exactly like one that covered it. So every evidence comment — and the committed `evidence/<ticket>/verdict.json` — carries two extra sections, defined in full by the `claim-evidence-mapping` rule:
157
+
158
+ - **Artifact identity** — what the evidence was collected against, as values rather than placeholders: the `repository`, the `head_sha` the run observed, the `environment`, and per artifact its `sha256` digest and `captured_at`. Defined in full by the `claim-evidence-mapping` rule.
159
+ - **Not established** — a **required, never-omitted** heading listing what the verification did *not* prove: boundaries not exercised, environments not tested, behavior consciously out of scope. When nothing is outstanding it still renders, reading `None outstanding — reviewed`. It is never blank.
160
+
161
+ The committed verdict carries the machine-readable half:
162
+
163
+ ```
164
+ evidence/<ticket>/verdict.json
165
+ not_established: [] # what was NOT proved; may be empty
166
+ not_established_reviewed: true # attests the list was reviewed; may NEVER be omitted
167
+ artifact: { repository, base_sha, head_sha, build_id, environment, observed_at }
168
+ evidence: [ { evidence_id, kind, locator,
169
+ artifact_head_sha, # the head_sha in force when THIS artifact was captured
170
+ sha256, # content digest of the committed evidence file
171
+ captured_at } ]
172
+ ```
173
+
174
+ `artifact.head_sha` pins the build the verification observed; each entry's `sha256` pins the bytes. An entry whose `artifact_head_sha` differs from `artifact.head_sha` is an `artifact_mismatch` and a recomputed digest that disagrees is an `evidence_digest_mismatch` — each fails loudly, naming both SHAs or the evidence id, and blocks completion. At completion the pinned `head_sha` is reconciled against **the merged head** using the ancestry + deploy-run definition of "what shipped" that `lisa-drive-pr-to-merge` already owns (cite it; there is no second definition): pre-merge evidence counts only when its head is a parent of the merge, and on a merge-race mismatch verification re-runs against the merged head before completion is declared.
175
+
176
+ The list may be empty; the flag may not be missing. An absent `not_established_reviewed` is indistinguishable from nobody having asked the question, so the evidence-posting gate in `tracker-evidence` refuses the post, and the Stop-hook gate reports it as a v2 contract violation (advisory until `verification.gate.enforceBoundaries` is ratcheted on). This generalizes the required, never-empty `Known limits` field of `lisa-improve-harness` to every evidence surface.
177
+
178
+ The boundary each artifact type reaches — and therefore which claim a captured artifact can discharge — is the `claim-evidence-mapping` rule's taxonomy; the type table above is its evidence-kind source.
179
+
154
180
  ### Cross-work-item evidence references are non-claiming
155
181
 
156
182
  When prose needs to point at evidence declared by another work item, use the dedicated reference form:
@@ -24,6 +24,16 @@ Upload captured evidence and generated templates to the GitHub PR description an
24
24
  - `comment.md` — GitHub markdown body for both the issue comment and the PR description's `## Evidence` section.
25
25
  - (Optional) `comment.txt` — kept for parity with the JIRA path; not used here.
26
26
 
27
+ ## Comment-body preflight (required)
28
+
29
+ Before posting or updating anything, check the evidence body (`comment.md`, and `comment.txt` where this skill uses it):
30
+
31
+ - It contains a `## Not established` heading. That heading is **never omitted and never blank** — when nothing is outstanding it still renders `None outstanding — reviewed`; otherwise it names, in plain operator language, what the verification did not prove.
32
+ - The accompanying verdict carries `not_established_reviewed: true` (the list may be empty; the flag may never be omitted).
33
+ - It contains a `## Artifact identity` heading carrying **values, not placeholders** — the repository, the `head_sha` the verification observed, the `environment`, and per artifact its `sha256` digest and `captured_at`. **Refuse to post** a body whose identity heading is absent or unpopulated, or whose recorded `artifact_head_sha` disagrees with the verdict's `artifact.head_sha` — report the evidence id and **both SHAs**. Definition: the `claim-evidence-mapping` rule.
34
+
35
+ If either is missing, **refuse to post**: stop and report the missing Not-established review to the caller instead of publishing. Composing the body is `lisa-tracker-evidence`'s job (see its UI Evidence Checklist); this skill only refuses to publish one that omits the section. The section is defined by the `claim-evidence-mapping` rule and generalizes `lisa-improve-harness`'s required, never-empty `Known limits` field.
36
+
27
37
  ## Workflow
28
38
 
29
39
  1. **Resolve refs**
@@ -42,6 +42,16 @@ Upload captured evidence and generated templates to GitHub PR description and JI
42
42
  - `comment.txt` — JIRA wiki markup (generated by `generate-templates.py`)
43
43
  - `comment.md` — GitHub markdown (generated by `generate-templates.py`)
44
44
 
45
+ ## Comment-body preflight (required)
46
+
47
+ Before posting or updating anything, check the evidence body (`comment.md`, and `comment.txt` where this skill uses it):
48
+
49
+ - It contains a `## Not established` heading. That heading is **never omitted and never blank** — when nothing is outstanding it still renders `None outstanding — reviewed`; otherwise it names, in plain operator language, what the verification did not prove.
50
+ - The accompanying verdict carries `not_established_reviewed: true` (the list may be empty; the flag may never be omitted).
51
+ - It contains a `## Artifact identity` heading carrying **values, not placeholders** — the repository, the `head_sha` the verification observed, the `environment`, and per artifact its `sha256` digest and `captured_at`. **Refuse to post** a body whose identity heading is absent or unpopulated, or whose recorded `artifact_head_sha` disagrees with the verdict's `artifact.head_sha` — report the evidence id and **both SHAs**. Definition: the `claim-evidence-mapping` rule.
52
+
53
+ If either is missing, **refuse to post**: stop and report the missing Not-established review to the caller instead of publishing. Composing the body is `lisa-tracker-evidence`'s job (see its UI Evidence Checklist); this skill only refuses to publish one that omits the section. The section is defined by the `claim-evidence-mapping` rule and generalizes `lisa-improve-harness`'s required, never-empty `Known limits` field.
54
+
45
55
  ## Usage
46
56
 
47
57
  ```bash
@@ -41,6 +41,16 @@ The caller must produce:
41
41
 
42
42
  If any of these are missing, stop and report.
43
43
 
44
+ ## Comment-body preflight (required)
45
+
46
+ Before posting or updating anything, check the evidence body (`comment.md`, and `comment.txt` where this skill uses it):
47
+
48
+ - It contains a `## Not established` heading. That heading is **never omitted and never blank** — when nothing is outstanding it still renders `None outstanding — reviewed`; otherwise it names, in plain operator language, what the verification did not prove.
49
+ - The accompanying verdict carries `not_established_reviewed: true` (the list may be empty; the flag may never be omitted).
50
+ - It contains a `## Artifact identity` heading carrying **values, not placeholders** — the repository, the `head_sha` the verification observed, the `environment`, and per artifact its `sha256` digest and `captured_at`. **Refuse to post** a body whose identity heading is absent or unpopulated, or whose recorded `artifact_head_sha` disagrees with the verdict's `artifact.head_sha` — report the evidence id and **both SHAs**. Definition: the `claim-evidence-mapping` rule.
51
+
52
+ If either is missing, **refuse to post**: stop and report the missing Not-established review to the caller instead of publishing. Composing the body is `lisa-tracker-evidence`'s job (see its UI Evidence Checklist); this skill only refuses to publish one that omits the section. The section is defined by the `claim-evidence-mapping` rule and generalizes `lisa-improve-harness`'s required, never-empty `Known limits` field.
53
+
44
54
  ## Phase 1 — Resolve Linear Issue
45
55
 
46
56
  1. Parse the identifier from `$ARGUMENTS`.