@codyswann/lisa 2.263.0 → 2.265.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/core/upstream-evidence-manifest.d.ts.map +1 -1
- package/dist/core/upstream-evidence-manifest.js +13 -10
- package/dist/core/upstream-evidence-manifest.js.map +1 -1
- package/package.json +1 -1
- package/plugins/lisa/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa/.codex-plugin/skills/lisa-github-evidence/SKILL.md +10 -0
- package/plugins/lisa/.codex-plugin/skills/lisa-jira-evidence/SKILL.md +10 -0
- package/plugins/lisa/.codex-plugin/skills/lisa-linear-evidence/SKILL.md +10 -0
- package/plugins/lisa/.codex-plugin/skills/lisa-security-review/SKILL.md +69 -2
- package/plugins/lisa/.codex-plugin/skills/lisa-security-zap-scan/SKILL.md +22 -2
- package/plugins/lisa/.codex-plugin/skills/lisa-tracker-evidence/SKILL.md +10 -3
- package/plugins/lisa/agents/security-specialist.md +19 -2
- package/plugins/lisa/rules/eager/claim-evidence-mapping.md +30 -3
- package/plugins/lisa/rules/reference/claim-evidence-mapping.md +155 -8
- package/plugins/lisa/rules/reference/verification.md +26 -0
- package/plugins/lisa/skills/lisa-github-evidence/SKILL.md +10 -0
- package/plugins/lisa/skills/lisa-jira-evidence/SKILL.md +10 -0
- package/plugins/lisa/skills/lisa-linear-evidence/SKILL.md +10 -0
- package/plugins/lisa/skills/lisa-security-review/SKILL.md +69 -2
- package/plugins/lisa/skills/lisa-security-zap-scan/SKILL.md +22 -2
- package/plugins/lisa/skills/lisa-tracker-evidence/SKILL.md +10 -3
- package/plugins/lisa-agy/agents/security-specialist.md +19 -2
- package/plugins/lisa-agy/plugin.json +1 -1
- package/plugins/lisa-agy/skills/lisa-github-evidence/SKILL.md +10 -0
- package/plugins/lisa-agy/skills/lisa-jira-evidence/SKILL.md +10 -0
- package/plugins/lisa-agy/skills/lisa-linear-evidence/SKILL.md +10 -0
- package/plugins/lisa-agy/skills/lisa-security-review/SKILL.md +69 -2
- package/plugins/lisa-agy/skills/lisa-security-zap-scan/SKILL.md +22 -2
- package/plugins/lisa-agy/skills/lisa-tracker-evidence/SKILL.md +10 -3
- package/plugins/lisa-cdk/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-cdk/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-cdk-agy/plugin.json +1 -1
- package/plugins/lisa-cdk-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-cdk-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-copilot/agents/security-specialist.agent.md +19 -2
- package/plugins/lisa-copilot/rules/eager/claim-evidence-mapping.md +30 -3
- package/plugins/lisa-copilot/rules/reference/claim-evidence-mapping.md +155 -8
- package/plugins/lisa-copilot/rules/reference/verification.md +26 -0
- package/plugins/lisa-copilot/skills/lisa-github-evidence/SKILL.md +10 -0
- package/plugins/lisa-copilot/skills/lisa-jira-evidence/SKILL.md +10 -0
- package/plugins/lisa-copilot/skills/lisa-linear-evidence/SKILL.md +10 -0
- package/plugins/lisa-copilot/skills/lisa-security-review/SKILL.md +69 -2
- package/plugins/lisa-copilot/skills/lisa-security-zap-scan/SKILL.md +22 -2
- package/plugins/lisa-copilot/skills/lisa-tracker-evidence/SKILL.md +10 -3
- package/plugins/lisa-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-cursor/agents/security-specialist.md +19 -2
- package/plugins/lisa-cursor/rules/claim-evidence-mapping-reference.mdc +155 -8
- package/plugins/lisa-cursor/rules/claim-evidence-mapping.mdc +30 -3
- package/plugins/lisa-cursor/rules/verification-reference.mdc +26 -0
- package/plugins/lisa-cursor/skills/lisa-github-evidence/SKILL.md +10 -0
- package/plugins/lisa-cursor/skills/lisa-jira-evidence/SKILL.md +10 -0
- package/plugins/lisa-cursor/skills/lisa-linear-evidence/SKILL.md +10 -0
- package/plugins/lisa-cursor/skills/lisa-security-review/SKILL.md +69 -2
- package/plugins/lisa-cursor/skills/lisa-security-zap-scan/SKILL.md +22 -2
- package/plugins/lisa-cursor/skills/lisa-tracker-evidence/SKILL.md +10 -3
- package/plugins/lisa-expo/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-expo/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-expo-agy/plugin.json +1 -1
- package/plugins/lisa-expo-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-expo-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-harper-fabric/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-harper-fabric/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-harper-fabric-agy/plugin.json +1 -1
- package/plugins/lisa-harper-fabric-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-harper-fabric-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-nestjs/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-nestjs/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-nestjs-agy/plugin.json +1 -1
- package/plugins/lisa-nestjs-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-nestjs-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-openclaw/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-openclaw/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-openclaw-agy/plugin.json +1 -1
- package/plugins/lisa-openclaw-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-openclaw-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-phaser/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-phaser/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-phaser-agy/plugin.json +1 -1
- package/plugins/lisa-phaser-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-phaser-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-rails/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-rails/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-rails-agy/plugin.json +1 -1
- package/plugins/lisa-rails-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-rails-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-typescript/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-typescript/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-typescript-agy/plugin.json +1 -1
- package/plugins/lisa-typescript-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-typescript-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-wiki/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-wiki/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-wiki-agy/plugin.json +1 -1
- package/plugins/lisa-wiki-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-wiki-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/src/base/agents/security-specialist.md +19 -2
- package/plugins/src/base/rules/eager/claim-evidence-mapping.md +30 -3
- package/plugins/src/base/rules/reference/claim-evidence-mapping.md +155 -8
- package/plugins/src/base/rules/reference/verification.md +26 -0
- package/plugins/src/base/skills/lisa-github-evidence/SKILL.md +10 -0
- package/plugins/src/base/skills/lisa-jira-evidence/SKILL.md +10 -0
- package/plugins/src/base/skills/lisa-linear-evidence/SKILL.md +10 -0
- package/plugins/src/base/skills/lisa-security-review/SKILL.md +69 -2
- package/plugins/src/base/skills/lisa-security-zap-scan/SKILL.md +22 -2
- package/plugins/src/base/skills/lisa-tracker-evidence/SKILL.md +10 -3
|
@@ -85,12 +85,11 @@ defines the names; it stores nothing:
|
|
|
85
85
|
| `required_evidence_kinds` | the evidence kind(s) that reach that boundary, from the `verification` artifact-type set |
|
|
86
86
|
|
|
87
87
|
A claim whose `required_evidence_kinds` has no captured, reaching artifact is **Not established** —
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
assume its section is present in this branch.
|
|
88
|
+
defined in full, with its evidence templates, in the section below. Artifact identity — what makes
|
|
89
|
+
two captured artifacts the same or different — is defined in the *Artifact identity* section below
|
|
90
|
+
(shipped by BCE-4, #1838). The bucket a security-sensitive claim lands in is defined in the *Security
|
|
91
|
+
buckets* section below (shipped by BCE-5, #1839). Where such a sibling surface is not installed in a
|
|
92
|
+
given branch, name what you can and continue.
|
|
94
93
|
|
|
95
94
|
### Worked example
|
|
96
95
|
|
|
@@ -119,13 +118,161 @@ Claim: "The service is deployed and healthy."
|
|
|
119
118
|
response from the target environment.
|
|
120
119
|
```
|
|
121
120
|
|
|
121
|
+
## Artifact identity — what the evidence was collected against
|
|
122
|
+
|
|
123
|
+
A claim reaches only as far as its evidence's *kind*; it applies only to the *artifact* that evidence
|
|
124
|
+
was collected against. **Artifact identity is what makes two captured artifacts the same or
|
|
125
|
+
different**: one repository, at one commit, in one environment, at one moment. Evidence that does not
|
|
126
|
+
say which artifact it observed is not evidence about anything in particular — it silently transfers
|
|
127
|
+
to whatever ships next, which is exactly how an auto-merge race ships code no verification ever
|
|
128
|
+
touched.
|
|
129
|
+
|
|
130
|
+
Identity is carried in two places, and they must agree:
|
|
131
|
+
|
|
132
|
+
| Where | Field | What it pins |
|
|
133
|
+
|---|---|---|
|
|
134
|
+
| `artifact` (once per verdict) | `repository` | the `owner/repo` the run observed |
|
|
135
|
+
| | `base_sha` | the base the change was measured against |
|
|
136
|
+
| | `head_sha` | **the commit the verification actually observed** — required for a v2 pass |
|
|
137
|
+
| | `build_id` | the build/run the evidence came from, where one exists |
|
|
138
|
+
| | `environment` | where it ran (local, preview, staging, production) |
|
|
139
|
+
| | `observed_at` | when the run observed it (ISO-8601 UTC) |
|
|
140
|
+
| `evidence[]` (per artifact) | `artifact_head_sha` | the `head_sha` in force **when that artifact was captured** |
|
|
141
|
+
| | `sha256` | content digest of the committed evidence file |
|
|
142
|
+
| | `captured_at` | when that artifact was captured (ISO-8601 UTC) |
|
|
143
|
+
|
|
144
|
+
`head_sha` pins the *build*; `sha256` pins the *bytes*. Together they answer both identity questions:
|
|
145
|
+
"which artifact was this collected against" and "is this still the artifact that was collected". The
|
|
146
|
+
`sha256` + commit-ref discipline is the same one Lisa's upstream-evidence manifest already uses — it
|
|
147
|
+
is cited as prior art here, not reinvented.
|
|
148
|
+
|
|
149
|
+
### Two identity failures, both loud
|
|
150
|
+
|
|
151
|
+
- **`artifact_mismatch`** — an `evidence[]` entry whose `artifact_head_sha` differs from the verdict's
|
|
152
|
+
`artifact.head_sha`. The evidence describes a different build than the one the verdict claims. The
|
|
153
|
+
identity check fails **loudly, naming both SHAs** (the evidence's and the verdict's) so an operator
|
|
154
|
+
can see which build each half is talking about.
|
|
155
|
+
- **`evidence_digest_mismatch`** — on read, recompute the `sha256` of each committed evidence file. If
|
|
156
|
+
the bytes no longer match the recorded digest, the artifact has changed since it was recorded: the
|
|
157
|
+
check fails, it names the evidence id, and blocks completion. A digest that cannot be recomputed
|
|
158
|
+
(file absent) is the same failure.
|
|
159
|
+
|
|
160
|
+
Neither failure is ever summarized as "verification failed". Name the field, the ids, and both SHAs —
|
|
161
|
+
a person who does not code reads this at the gate (`factory-model` rule 5). Both checks are
|
|
162
|
+
**advisory-first**: reported to the operator but non-blocking until
|
|
163
|
+
`verification.gate.enforceBoundaries` is `true` in `.lisa.config.json` — the same ratchet flag the
|
|
164
|
+
boundary checks and the Stop-hook gate ride.
|
|
165
|
+
|
|
166
|
+
### The merge race — one definition, two guards
|
|
167
|
+
|
|
168
|
+
"The artifact that shipped" is already defined once, in `lisa-drive-pr-to-merge`: the **ancestry**
|
|
169
|
+
check (is the verified commit an ancestor of the merged base branch, and is the merge commit's parent
|
|
170
|
+
that verified head rather than a stale one) **plus** the **deploy-run** check (a deploy/release run
|
|
171
|
+
actually fired for the merge SHA or an including descendant). Cite that definition; never write a
|
|
172
|
+
second one. Identity reconciliation is the same definition read from the evidence side:
|
|
173
|
+
|
|
174
|
+
- Evidence collected on a **pre-merge head** is valid for the merge commit **only when that head is a
|
|
175
|
+
parent of the merge** — i.e. it satisfies that skill's ancestry check against the merged head. If a
|
|
176
|
+
late commit raced past the merge, the merged head is not the verified one and the evidence does not
|
|
177
|
+
transfer.
|
|
178
|
+
- Ancestry alone is never enough. That skill's rule stands unchanged here: **never report shipped on
|
|
179
|
+
ancestry alone** — the deploy-run check must also pass before completion is declared.
|
|
180
|
+
- On mismatch (`artifact.head_sha` ≠ the reconciled merge SHA), completion is not declared. Flag the
|
|
181
|
+
mismatch naming both SHAs and **re-run verification against the merged head**; the re-run's verdict
|
|
182
|
+
is what may declare completion.
|
|
183
|
+
|
|
184
|
+
Two guards, one definition — so they cannot disagree about what shipped.
|
|
185
|
+
|
|
186
|
+
## Security buckets — the impact-or-exploitability bar
|
|
187
|
+
|
|
188
|
+
A security-shaped claim is a claim like any other: it reaches only as far as its evidence. Applied to
|
|
189
|
+
security findings (shipped by BCE-5, #1839), that yields two buckets and one bar:
|
|
190
|
+
|
|
191
|
+
- **Security (proven)** — the finding carries **both** a `reproducer` (an evidence ref of a kind that
|
|
192
|
+
reaches the claim's boundary, per the taxonomy above — e.g. an `http-transcript` for an injection
|
|
193
|
+
claim at `http-api`) **and** a bounded `impact`/exploitability statement.
|
|
194
|
+
- **Security (unproven)** — **missing either**. It keeps a one-line `reason` and **stays in the
|
|
195
|
+
security section**: a reproducer-less finding is **never auto-demoted** to a `maintenance` bucket,
|
|
196
|
+
because under-reporting a real vulnerability is the worse failure.
|
|
197
|
+
|
|
198
|
+
The unproven bucket's label is the single configurable policy point —
|
|
199
|
+
`security.review.unprovenBucket` in `.lisa.config.json`, default `security-unproven`. An owner who
|
|
200
|
+
prefers true demotion flips that one value; no other classification logic changes. Like the boundary
|
|
201
|
+
checks, bucketing is **advisory** — it shapes the report, not the merge — until
|
|
202
|
+
`verification.gate.enforceBoundaries` is `true`.
|
|
203
|
+
|
|
204
|
+
The full procedure — per-finding fields, report shape, ZAP alignment — lives in the
|
|
205
|
+
`lisa-security-review` skill (with `lisa-security-zap-scan` citing it); dependency-CVE remediation
|
|
206
|
+
keeps its separate ladder in `security-audit-handling`. Cite those slugs; do not restate them here.
|
|
207
|
+
|
|
208
|
+
## The "Not established" section — required, never omitted
|
|
209
|
+
|
|
210
|
+
Every report that asserts something was verified also states, in the same breath, **what it did not
|
|
211
|
+
establish**. This is one section, it is **required**, and it is **never omitted and never blank** —
|
|
212
|
+
on the evidence comment, in the committed `evidence/<ticket>/verdict.json`, and in the
|
|
213
|
+
`verification-status.json` verdict BCE-2's gate reads.
|
|
214
|
+
|
|
215
|
+
**Where it appears and what it says**
|
|
216
|
+
|
|
217
|
+
| Surface | Shape |
|
|
218
|
+
|---|---|
|
|
219
|
+
| Evidence comment (tracker + PR `## Evidence` section) | a `## Not established` heading, one plain-language bullet per item |
|
|
220
|
+
| `evidence/<ticket>/verdict.json` | `not_established: []` plus `not_established_reviewed: true` |
|
|
221
|
+
| `.lisa/verification-status.json` (schema v2) | per-claim `not_established[]` plus the top-level `not_established_reviewed` flag |
|
|
222
|
+
|
|
223
|
+
**The empty case is not the omitted case.** A verification that genuinely left nothing unproved still
|
|
224
|
+
renders the heading, with the single line:
|
|
225
|
+
|
|
226
|
+
```text
|
|
227
|
+
## Not established
|
|
228
|
+
|
|
229
|
+
None outstanding — reviewed
|
|
230
|
+
```
|
|
231
|
+
|
|
232
|
+
An absent heading, or a heading with nothing under it, is a defect — it is indistinguishable from
|
|
233
|
+
never having asked the question. The list may be empty; the section may not be blank.
|
|
234
|
+
|
|
235
|
+
**Machine-readable semantics (as shipped by BCE-2, #1836).** `not_established` is the list — it may
|
|
236
|
+
be empty. `not_established_reviewed` is the boolean attestation that the list was actually
|
|
237
|
+
reviewed — **the flag may never be omitted**. That asymmetry is the whole mechanism: an empty list
|
|
238
|
+
plus a present flag means "we looked and found nothing outstanding"; an absent flag means nobody
|
|
239
|
+
looked. The Stop-hook gate treats an absent flag as a v2 contract violation, reported to stderr and
|
|
240
|
+
**advisory** until `verification.gate.enforceBoundaries` is `true` in `.lisa.config.json` (the same
|
|
241
|
+
ratchet flag the boundary checks ride); evidence surfaces refuse the post on the same terms.
|
|
242
|
+
|
|
243
|
+
**What belongs under the heading** — written in operator voice (`factory-model` rule 5: a person who
|
|
244
|
+
does not code reads this at the gate), not in engineering shorthand:
|
|
245
|
+
|
|
246
|
+
- **Boundaries not exercised** — a claim's boundary that no captured artifact reached. *"The checkout
|
|
247
|
+
button was proved in the browser; the order's persisted row was never queried, so the `data`
|
|
248
|
+
boundary is not established."*
|
|
249
|
+
- **Environments not tested** — where it was and was not run. *"Checked on production Chrome at
|
|
250
|
+
1440×900 only. Not checked on mobile Safari, and not checked against the staging database."*
|
|
251
|
+
- **Claims consciously out of scope** — deliberately excluded behavior, named so nobody infers it.
|
|
252
|
+
*"Refunds were not touched or tested; this change covers new orders only."*
|
|
253
|
+
- **Anything a green quality check might be mistaken for proving.** *"Unit tests pass for the submit
|
|
254
|
+
handler. That establishes the code-unit boundary only — it is not evidence the button works."*
|
|
255
|
+
|
|
256
|
+
Each item names the thing, not a category: "not tested on mobile Safari" is usable at a gate; "some
|
|
257
|
+
environments untested" is not.
|
|
258
|
+
|
|
259
|
+
### Philosophical precedent for this section
|
|
260
|
+
|
|
261
|
+
This generalizes `lisa-improve-harness`'s **`Known limits`** field — a required, never-empty line on
|
|
262
|
+
every result record. That skill says it plainly:
|
|
263
|
+
*"A record with nothing in it is invalid on its face"* — because a single-trajectory loop always has
|
|
264
|
+
limits. The same is true of any verification:
|
|
265
|
+
it ran somewhere, on something, once. `Known limits` (one record) and `not_established` (every claim
|
|
266
|
+
in the factory) are the same discipline with the same never-empty rule; read either and you should
|
|
267
|
+
recognize the other.
|
|
268
|
+
|
|
122
269
|
## Philosophical precedent
|
|
123
270
|
|
|
124
271
|
This generalizes the **bounded-claim discipline** of `lisa-improve-harness`: one trajectory supports
|
|
125
272
|
one trajectory's claim, and a result record may claim only what its cited evidence reaches. Here the
|
|
126
273
|
same discipline is applied to every claim in the factory — a claim reaches exactly as far as the
|
|
127
|
-
*kind* of evidence behind it, and no further.
|
|
128
|
-
|
|
274
|
+
*kind* of evidence behind it, and no further. The *Not established* half of that discipline is a
|
|
275
|
+
first-class report state — see the required, never-omitted section above.
|
|
129
276
|
|
|
130
277
|
## No behavior change; degrade, never block
|
|
131
278
|
|
|
@@ -42,9 +42,36 @@ rule 5).
|
|
|
42
42
|
A claim carries three fields — `claim_id`, `boundary`, and `required_evidence_kinds` — named here so
|
|
43
43
|
every downstream surface uses one spelling. This ticket only writes the contract down; the schema and
|
|
44
44
|
gate that make these fields executable ship with **BCE-2 (#1836)** — do not assume that surface is
|
|
45
|
-
present in this branch.
|
|
46
|
-
|
|
47
|
-
|
|
45
|
+
present in this branch.
|
|
46
|
+
|
|
47
|
+
## Security buckets (conservative by default)
|
|
48
|
+
|
|
49
|
+
A security finding is **proven** only with both a reproducer of a reaching kind and a bounded
|
|
50
|
+
impact/exploitability statement; missing either, it renders **unproven** with its reason and **stays
|
|
51
|
+
in the security section** — never auto-demoted to maintenance. The label is one policy point,
|
|
52
|
+
`security.review.unprovenBucket` (default `security-unproven`); the procedure lives in the
|
|
53
|
+
`lisa-security-review` skill.
|
|
54
|
+
|
|
55
|
+
## Artifact identity (pinned, never assumed)
|
|
56
|
+
|
|
57
|
+
A claim applies only to the artifact its evidence was collected against. Every verdict pins
|
|
58
|
+
`artifact.head_sha` — the commit the run observed — and every evidence entry pins the
|
|
59
|
+
`artifact_head_sha` in force when it was captured, its `sha256` content digest, and `captured_at`.
|
|
60
|
+
Evidence collected on a pre-merge head is valid for the merge commit **only** when that head is a
|
|
61
|
+
parent of the merge, per the ancestry + deploy-run definition of "what shipped" that
|
|
62
|
+
`lisa-drive-pr-to-merge` already owns — cite it, never write a second one. A mismatched SHA or a
|
|
63
|
+
recomputed digest that disagrees fails loudly, naming both SHAs / the evidence id, and verification
|
|
64
|
+
re-runs against the merged head before completion is declared. Full definition: the reference body.
|
|
65
|
+
|
|
66
|
+
## Not established (required, never omitted)
|
|
67
|
+
|
|
68
|
+
A claim with no reaching evidence is **Not established**, and every report says so out loud. Each
|
|
69
|
+
evidence comment and verdict carries a `Not established` section listing what was *not* proved —
|
|
70
|
+
boundaries not exercised, environments not tested, behavior consciously out of scope. It is never
|
|
71
|
+
omitted and never blank: with nothing outstanding it still renders `None outstanding — reviewed`, and
|
|
72
|
+
`not_established_reviewed` attests the list was reviewed even when the list itself is empty. This
|
|
73
|
+
generalizes the required, never-empty `Known limits` field of `lisa-improve-harness`. Full definition
|
|
74
|
+
and operator-voice exemplars: the reference body.
|
|
48
75
|
|
|
49
76
|
## No behavior change; degrade, never block
|
|
50
77
|
|
|
@@ -156,6 +156,32 @@ Do not invent types inline; if none fits, propose extending this table. The lega
|
|
|
156
156
|
|
|
157
157
|
The manifest is the single source of truth for "what evidence is required": authored once in the Validation Journey, enforced at write time, replayed during `tracker-journey` (which captures each artifact **in its declared type**), and checked again before the ticket closes. There is no second list to keep in sync.
|
|
158
158
|
|
|
159
|
+
### Every evidence surface names what it did NOT establish
|
|
160
|
+
|
|
161
|
+
An evidence comment that lists only what passed is unreadable at a gate: a journey that skipped an edge state looks exactly like one that covered it. So every evidence comment — and the committed `evidence/<ticket>/verdict.json` — carries two extra sections, defined in full by the `claim-evidence-mapping` rule:
|
|
162
|
+
|
|
163
|
+
- **Artifact identity** — what the evidence was collected against, as values rather than placeholders: the `repository`, the `head_sha` the run observed, the `environment`, and per artifact its `sha256` digest and `captured_at`. Defined in full by the `claim-evidence-mapping` rule.
|
|
164
|
+
- **Not established** — a **required, never-omitted** heading listing what the verification did *not* prove: boundaries not exercised, environments not tested, behavior consciously out of scope. When nothing is outstanding it still renders, reading `None outstanding — reviewed`. It is never blank.
|
|
165
|
+
|
|
166
|
+
The committed verdict carries the machine-readable half:
|
|
167
|
+
|
|
168
|
+
```
|
|
169
|
+
evidence/<ticket>/verdict.json
|
|
170
|
+
not_established: [] # what was NOT proved; may be empty
|
|
171
|
+
not_established_reviewed: true # attests the list was reviewed; may NEVER be omitted
|
|
172
|
+
artifact: { repository, base_sha, head_sha, build_id, environment, observed_at }
|
|
173
|
+
evidence: [ { evidence_id, kind, locator,
|
|
174
|
+
artifact_head_sha, # the head_sha in force when THIS artifact was captured
|
|
175
|
+
sha256, # content digest of the committed evidence file
|
|
176
|
+
captured_at } ]
|
|
177
|
+
```
|
|
178
|
+
|
|
179
|
+
`artifact.head_sha` pins the build the verification observed; each entry's `sha256` pins the bytes. An entry whose `artifact_head_sha` differs from `artifact.head_sha` is an `artifact_mismatch` and a recomputed digest that disagrees is an `evidence_digest_mismatch` — each fails loudly, naming both SHAs or the evidence id, and blocks completion. At completion the pinned `head_sha` is reconciled against **the merged head** using the ancestry + deploy-run definition of "what shipped" that `lisa-drive-pr-to-merge` already owns (cite it; there is no second definition): pre-merge evidence counts only when its head is a parent of the merge, and on a merge-race mismatch verification re-runs against the merged head before completion is declared.
|
|
180
|
+
|
|
181
|
+
The list may be empty; the flag may not be missing. An absent `not_established_reviewed` is indistinguishable from nobody having asked the question, so the evidence-posting gate in `tracker-evidence` refuses the post, and the Stop-hook gate reports it as a v2 contract violation (advisory until `verification.gate.enforceBoundaries` is ratcheted on). This generalizes the required, never-empty `Known limits` field of `lisa-improve-harness` to every evidence surface.
|
|
182
|
+
|
|
183
|
+
The boundary each artifact type reaches — and therefore which claim a captured artifact can discharge — is the `claim-evidence-mapping` rule's taxonomy; the type table above is its evidence-kind source.
|
|
184
|
+
|
|
159
185
|
### Cross-work-item evidence references are non-claiming
|
|
160
186
|
|
|
161
187
|
When prose needs to point at evidence declared by another work item, use the dedicated reference form:
|
|
@@ -24,6 +24,16 @@ Upload captured evidence and generated templates to the GitHub PR description an
|
|
|
24
24
|
- `comment.md` — GitHub markdown body for both the issue comment and the PR description's `## Evidence` section.
|
|
25
25
|
- (Optional) `comment.txt` — kept for parity with the JIRA path; not used here.
|
|
26
26
|
|
|
27
|
+
## Comment-body preflight (required)
|
|
28
|
+
|
|
29
|
+
Before posting or updating anything, check the evidence body (`comment.md`, and `comment.txt` where this skill uses it):
|
|
30
|
+
|
|
31
|
+
- It contains a `## Not established` heading. That heading is **never omitted and never blank** — when nothing is outstanding it still renders `None outstanding — reviewed`; otherwise it names, in plain operator language, what the verification did not prove.
|
|
32
|
+
- The accompanying verdict carries `not_established_reviewed: true` (the list may be empty; the flag may never be omitted).
|
|
33
|
+
- It contains a `## Artifact identity` heading carrying **values, not placeholders** — the repository, the `head_sha` the verification observed, the `environment`, and per artifact its `sha256` digest and `captured_at`. **Refuse to post** a body whose identity heading is absent or unpopulated, or whose recorded `artifact_head_sha` disagrees with the verdict's `artifact.head_sha` — report the evidence id and **both SHAs**. Definition: the `claim-evidence-mapping` rule.
|
|
34
|
+
|
|
35
|
+
If either is missing, **refuse to post**: stop and report the missing Not-established review to the caller instead of publishing. Composing the body is `lisa-tracker-evidence`'s job (see its UI Evidence Checklist); this skill only refuses to publish one that omits the section. The section is defined by the `claim-evidence-mapping` rule and generalizes `lisa-improve-harness`'s required, never-empty `Known limits` field.
|
|
36
|
+
|
|
27
37
|
## Workflow
|
|
28
38
|
|
|
29
39
|
1. **Resolve refs**
|
|
@@ -42,6 +42,16 @@ Upload captured evidence and generated templates to GitHub PR description and JI
|
|
|
42
42
|
- `comment.txt` — JIRA wiki markup (generated by `generate-templates.py`)
|
|
43
43
|
- `comment.md` — GitHub markdown (generated by `generate-templates.py`)
|
|
44
44
|
|
|
45
|
+
## Comment-body preflight (required)
|
|
46
|
+
|
|
47
|
+
Before posting or updating anything, check the evidence body (`comment.md`, and `comment.txt` where this skill uses it):
|
|
48
|
+
|
|
49
|
+
- It contains a `## Not established` heading. That heading is **never omitted and never blank** — when nothing is outstanding it still renders `None outstanding — reviewed`; otherwise it names, in plain operator language, what the verification did not prove.
|
|
50
|
+
- The accompanying verdict carries `not_established_reviewed: true` (the list may be empty; the flag may never be omitted).
|
|
51
|
+
- It contains a `## Artifact identity` heading carrying **values, not placeholders** — the repository, the `head_sha` the verification observed, the `environment`, and per artifact its `sha256` digest and `captured_at`. **Refuse to post** a body whose identity heading is absent or unpopulated, or whose recorded `artifact_head_sha` disagrees with the verdict's `artifact.head_sha` — report the evidence id and **both SHAs**. Definition: the `claim-evidence-mapping` rule.
|
|
52
|
+
|
|
53
|
+
If either is missing, **refuse to post**: stop and report the missing Not-established review to the caller instead of publishing. Composing the body is `lisa-tracker-evidence`'s job (see its UI Evidence Checklist); this skill only refuses to publish one that omits the section. The section is defined by the `claim-evidence-mapping` rule and generalizes `lisa-improve-harness`'s required, never-empty `Known limits` field.
|
|
54
|
+
|
|
45
55
|
## Usage
|
|
46
56
|
|
|
47
57
|
```bash
|
|
@@ -41,6 +41,16 @@ The caller must produce:
|
|
|
41
41
|
|
|
42
42
|
If any of these are missing, stop and report.
|
|
43
43
|
|
|
44
|
+
## Comment-body preflight (required)
|
|
45
|
+
|
|
46
|
+
Before posting or updating anything, check the evidence body (`comment.md`, and `comment.txt` where this skill uses it):
|
|
47
|
+
|
|
48
|
+
- It contains a `## Not established` heading. That heading is **never omitted and never blank** — when nothing is outstanding it still renders `None outstanding — reviewed`; otherwise it names, in plain operator language, what the verification did not prove.
|
|
49
|
+
- The accompanying verdict carries `not_established_reviewed: true` (the list may be empty; the flag may never be omitted).
|
|
50
|
+
- It contains a `## Artifact identity` heading carrying **values, not placeholders** — the repository, the `head_sha` the verification observed, the `environment`, and per artifact its `sha256` digest and `captured_at`. **Refuse to post** a body whose identity heading is absent or unpopulated, or whose recorded `artifact_head_sha` disagrees with the verdict's `artifact.head_sha` — report the evidence id and **both SHAs**. Definition: the `claim-evidence-mapping` rule.
|
|
51
|
+
|
|
52
|
+
If either is missing, **refuse to post**: stop and report the missing Not-established review to the caller instead of publishing. Composing the body is `lisa-tracker-evidence`'s job (see its UI Evidence Checklist); this skill only refuses to publish one that omits the section. The section is defined by the `claim-evidence-mapping` rule and generalizes `lisa-improve-harness`'s required, never-empty `Known limits` field.
|
|
53
|
+
|
|
44
54
|
## Phase 1 — Resolve Linear Issue
|
|
45
55
|
|
|
46
56
|
1. Parse the identifier from `$ARGUMENTS`.
|
|
@@ -16,6 +16,57 @@ Identify vulnerabilities, evaluate threats, and recommend mitigations for code c
|
|
|
16
16
|
5. **Check auth/authz** -- are access controls properly enforced for new endpoints or features?
|
|
17
17
|
6. **Review dependencies** -- do new dependencies introduce known vulnerabilities?
|
|
18
18
|
|
|
19
|
+
## The impact-or-exploitability bar
|
|
20
|
+
|
|
21
|
+
Severity is **earned, not pattern-matched**. Every security-shaped finding is classified
|
|
22
|
+
mechanically, before it is written up:
|
|
23
|
+
|
|
24
|
+
| Field | What it holds |
|
|
25
|
+
|-------|---------------|
|
|
26
|
+
| `reproducer` | an evidence ref of a kind that reaches the claim's boundary, or `none` |
|
|
27
|
+
| `impact` | a bounded impact/exploitability statement (who can do what, to what data, under what preconditions), or `unproven` |
|
|
28
|
+
| `reason` | one line saying why the finding landed in its bucket |
|
|
29
|
+
|
|
30
|
+
**The bar:** a finding is **proven** only when it carries **both** a reproducer **and** a bounded
|
|
31
|
+
impact statement. **Missing either ⇒ unproven.** No other input changes the bucket.
|
|
32
|
+
|
|
33
|
+
What counts as a reaching reproducer is defined by the `claim-evidence-mapping` contract (**BCE-1**,
|
|
34
|
+
#1835), not here: an injection claim at the `http-api` boundary needs an `http-transcript`; a UI
|
|
35
|
+
claim needs a `screenshot` or `recording`. A passing unit `test-run-log` reaches `code-unit` only and
|
|
36
|
+
never discharges either.
|
|
37
|
+
|
|
38
|
+
**Each field stands on its own.** The two halves are recorded independently: a finding with a bounded
|
|
39
|
+
impact but no reproducer keeps its impact statement verbatim and only `reproducer` reads `none`; a
|
|
40
|
+
finding with a reproducer but no bounded impact keeps the evidence ref and only `impact` reads
|
|
41
|
+
`unproven`. **Never overwrite** a field you actually have with a missing-value placeholder — the
|
|
42
|
+
`reason` line names which half is missing, and the surviving half is the head start the next reviewer
|
|
43
|
+
needs.
|
|
44
|
+
|
|
45
|
+
## The two buckets — conservative by default
|
|
46
|
+
|
|
47
|
+
Findings render in two clearly-labeled buckets: **Security (proven)** and **Security (unproven)**.
|
|
48
|
+
|
|
49
|
+
A reproducer-less finding **stays in the security section**, labeled `unproven` with its reason. It
|
|
50
|
+
is **never auto-demoted** to a `maintenance` bucket and it is **not removed** from the report —
|
|
51
|
+
under-reporting a real vulnerability is the worse failure, so the conservative default keeps it
|
|
52
|
+
visible where a security reader looks.
|
|
53
|
+
|
|
54
|
+
**Single policy point.** The unproven bucket's label is the only thing an owner may change:
|
|
55
|
+
`security.review.unprovenBucket` in `.lisa.config.json`, default `security-unproven`. An owner who
|
|
56
|
+
prefers true demotion sets it to a maintenance label; the finding then renders under that bucket and
|
|
57
|
+
**no other classification logic changes** — the bar, the fields, and the reasons are identical.
|
|
58
|
+
|
|
59
|
+
Write both buckets in operator voice (`factory-model` rule 5): a person who does not code reads this
|
|
60
|
+
at the gate. "Anyone who can reach the search box can read other customers' orders — reproduced with
|
|
61
|
+
the request transcript below" is usable; "possible SQLi in handler" is not.
|
|
62
|
+
|
|
63
|
+
This bar governs *code-review* security findings. Dependency CVE remediation keeps its own decision
|
|
64
|
+
ladder in the `security-audit-handling` rule — cite it, do not restate or fork it.
|
|
65
|
+
|
|
66
|
+
Bucketing is **advisory** — it shapes the report, it does not block a merge — on the same terms as
|
|
67
|
+
the boundary checks, which stay reporting-only until `verification.gate.enforceBoundaries` is `true`
|
|
68
|
+
in `.lisa.config.json`.
|
|
69
|
+
|
|
19
70
|
## Output Format
|
|
20
71
|
|
|
21
72
|
Structure findings as:
|
|
@@ -41,17 +92,33 @@ Structure findings as:
|
|
|
41
92
|
- [ ] No XSS vectors in user-facing output
|
|
42
93
|
- [ ] Dependencies free of known CVEs
|
|
43
94
|
|
|
44
|
-
###
|
|
45
|
-
- [
|
|
95
|
+
### Security (proven)
|
|
96
|
+
- [finding] -- where in the code, how to prevent
|
|
97
|
+
- reproducer: [evidence ref, e.g. evidence/<ticket>/http-transcript-01.txt]
|
|
98
|
+
- impact: [who can do what, to what data, under what preconditions]
|
|
99
|
+
- reason: reproducer + bounded impact
|
|
100
|
+
|
|
101
|
+
### Security (unproven)
|
|
102
|
+
- [finding] -- where in the code, how to prevent
|
|
103
|
+
- reproducer: [evidence ref if one exists, else `none`]
|
|
104
|
+
- impact: [bounded statement if one exists, else `unproven`]
|
|
105
|
+
- reason: [which half is missing -- e.g. "impact bounded, but never reproduced"]
|
|
106
|
+
-- kept in the security section, not demoted
|
|
46
107
|
|
|
47
108
|
### Recommendations
|
|
48
109
|
- [recommendation] -- priority (critical/warning/suggestion)
|
|
49
110
|
```
|
|
50
111
|
|
|
112
|
+
Rename the unproven heading only when `security.review.unprovenBucket` is set to something other
|
|
113
|
+
than `security-unproven`; everything else stays as written.
|
|
114
|
+
|
|
51
115
|
## Rules
|
|
52
116
|
|
|
53
117
|
- Focus on the specific changes proposed, not a full security audit of the entire codebase
|
|
54
118
|
- Flag only real risks -- do not invent hypothetical threats for internal tooling with no user input
|
|
119
|
+
- Classify every finding against the bar before writing it up; never leave a finding unbucketed
|
|
120
|
+
- Never silently drop or downgrade a finding out of the security section -- `unproven` is the
|
|
121
|
+
conservative landing spot, and the reason line says why
|
|
55
122
|
- Prioritize OWASP Top 10 vulnerabilities
|
|
56
123
|
- If the changes are purely internal (config, refactoring, docs), report "No security concerns" and explain why
|
|
57
124
|
- Always check `.gitleaksignore` patterns to understand what secrets scanning is already in place
|
|
@@ -22,10 +22,30 @@ Run a ZAP baseline security scan against the local application.
|
|
|
22
22
|
- After the scan completes, read `zap-report.html` (or `zap-report.md` for text)
|
|
23
23
|
- Summarize findings:
|
|
24
24
|
- Total number of alerts by risk level (High, Medium, Low, Informational)
|
|
25
|
-
-
|
|
25
|
+
- **Every alert reaches classification** -- High, Medium, Low, and Informational alike. Risk
|
|
26
|
+
level orders the summary; it never filters it. **Nothing is dropped before classification**,
|
|
27
|
+
so no alert can leave the report unclassified. Medium+ alerts are listed first, in full (rule
|
|
28
|
+
ID, name, recommended fix); Low/Informational alerts are still listed, bucketed, and given a
|
|
29
|
+
`reason`, even when compressed to one line each.
|
|
26
30
|
- Categorize findings as "infrastructure-level" (fix at CDN/proxy) vs "application-level" (fix in code)
|
|
27
31
|
|
|
28
|
-
4. **
|
|
32
|
+
4. **Apply the impact-or-exploitability bar** -- the same bar the `lisa-security-review` skill
|
|
33
|
+
defines; follow that skill, do not restate it. A ZAP alert is not a reproducer by itself: the
|
|
34
|
+
alert names a pattern, not an exercised impact path.
|
|
35
|
+
- **Security (proven)** -- the alert carries a reproducer **and** a bounded impact statement. The
|
|
36
|
+
reproducer counts only if its evidence kind **reaches the claim's boundary** under the
|
|
37
|
+
`claim-evidence-mapping` contract (BCE-1, #1835): a ZAP request/response transcript is an
|
|
38
|
+
`http-transcript` and reaches the `http-api` boundary only. An alert whose claim is about
|
|
39
|
+
rendered UI (`browser`) or persisted state (`data`) needs evidence at *that* boundary -- a
|
|
40
|
+
transcript never proves it.
|
|
41
|
+
- **Security (unproven)** -- everything else, each with a one-line `reason` (typically
|
|
42
|
+
"alert only, no reproducer / no bounded impact", or "transcript does not reach the claim's
|
|
43
|
+
boundary"). Unproven alerts are **not dropped** and not demoted out of the security summary --
|
|
44
|
+
they render in the unproven bucket so a reader still sees them.
|
|
45
|
+
- Rename the unproven heading only if `security.review.unprovenBucket` is set to something other
|
|
46
|
+
than `security-unproven`; no other classification changes.
|
|
47
|
+
|
|
48
|
+
5. **Handle failures**:
|
|
29
49
|
- If the scan failed, explain what failed and suggest concrete remediation steps
|
|
30
50
|
|
|
31
51
|
## Execution
|
|
@@ -28,6 +28,10 @@ See the `config-resolution` rule for configuration and dispatch table.
|
|
|
28
28
|
- Never post evidence to a different ticket than the one named — `$ARGUMENTS` is the source of truth.
|
|
29
29
|
- Never invent a verify-specific usage footer. Evidence artifact usage must flow through `lisa-usage-accounting`, preserve the canonical `## Lisa Usage` section, and surface `source: unavailable` explicitly when the runtime cannot provide trustworthy numbers.
|
|
30
30
|
- **Evidence-manifest gate (leaf work units).** Before dispatching to a vendor skill that transitions the ticket, confirm `EVIDENCE_DIR` contains a non-empty artifact **of the declared type** for every typed `[EVIDENCE: <artifact-type>: <name>]` marker declared in the ticket's Validation Journey — a `screenshot` marker needs an actual image, an `http-transcript` marker needs the request + response text, a `perf-trace` marker needs measured numbers; a prose claim satisfies nothing. If any declared marker has no captured artifact, an empty one, or one whose content/extension does not match its declared type, stop and report the offending markers by name instead of posting — a leaf work unit (Bug / Task / Sub-task / Improvement) may not advance to its review/Done state with an unsatisfied manifest (see the "Per-Work-Unit Evidence Contract" in the `verification` rule). Epics / Stories / Spikes, and leaf units without a Validation Journey, are exempt.
|
|
31
|
+
- **Claim↔boundary binding (S14 upgrade).** Satisfying the manifest by *type* is not enough: each `[EVIDENCE: <artifact-type>: <name>]` marker is also cited for a **claim**, and the marker's artifact type must **reach the boundary that claim declares** per the `claim-evidence-mapping` rule's taxonomy. A `browser` claim (user-visible UI behavior) is reached by `screenshot` / `recording` and never by a unit `test-run-log`; an `http-api` claim needs an `http-transcript`; a `deploy-health` claim needs a `deploy-log` and no pre-deploy artifact. On a mismatch, report the offending marker **by name** together with the claim, its boundary, and the **required evidence kinds** for that boundary — e.g. `[EVIDENCE: test-run-log: unit-suite] cited for claim AC-2 [boundary browser] — required evidence kinds: screenshot, recording`. Swapping in a marker of a reaching kind with a real captured artifact satisfies the gate. **Advisory-first:** until `verification.gate.enforceBoundaries` is `true` in `.lisa.config.json` (the same ratchet flag the Stop-hook gate reads, default `false`), a boundary mismatch is reported to the operator but does not block the post; once ratcheted on it refuses the post exactly like a missing or wrong-type artifact. Missing/empty/wrong-type artifacts keep refusing the post regardless of the flag.
|
|
32
|
+
- **The "Not established" section is required.** Before posting, confirm `evidence/comment.md` (and `comment.txt` where the vendor uses it) contains a `## Not established` heading and that the verdict carries `not_established_reviewed: true`. The heading is **never omitted and never blank**: with nothing outstanding it still renders `None outstanding — reviewed`; otherwise it lists, in plain operator language, each thing the verification did not prove — boundaries not exercised, environments not tested, behavior consciously out of scope. A comment with no such heading, a heading with nothing under it, or a verdict whose `not_established_reviewed` flag is absent is **refused**: stop and report the missing Not-established review instead of posting. (The list may be empty; the flag may never be omitted.) Definition and exemplars live in the `claim-evidence-mapping` rule; this generalizes `lisa-improve-harness`'s required, never-empty `Known limits` field.
|
|
33
|
+
- **Artifact identity is required, with values (S14 extension).** Before posting, confirm the comment body carries a `## Artifact identity` heading populated with **values, not placeholders**: the `repository`, the `head_sha` the verification observed, the `environment`, and for each committed artifact its `sha256` digest and `captured_at`. Then check both identity failures. **`artifact_mismatch`** — an evidence entry whose recorded `artifact_head_sha` differs from the verdict's `artifact.head_sha`: refuse the post and report the offending evidence id together with **both SHAs** (the one the artifact was captured at and the one the verdict claims), e.g. `EV-2 captured at 4f1c9ab but artifact.head_sha is 9de0c31`. **`evidence_digest_mismatch`** — recompute the `sha256` of each committed evidence file on read; if the bytes no longer match the recorded digest (or the file is absent), stop and report the evidence id by name. Never summarize either as "verification failed" — name the field, the ids, and both SHAs so a non-engineer can read it at the gate. **Advisory-first:** until `verification.gate.enforceBoundaries` is `true` in `.lisa.config.json`, an identity failure is reported to the operator but does not block the post; once ratcheted on it refuses the post exactly like a missing artifact. Definition: the `claim-evidence-mapping` rule.
|
|
34
|
+
- **Merge-race reconciliation (at completion, not at post).** The evidence's pinned `head_sha` is reconciled against the merged head using the ancestry + deploy-run definition of "what shipped" that `lisa-drive-pr-to-merge` already owns — cite that skill, never re-implement or restate its checks here. Pre-merge evidence counts for the merge commit only when its head is a parent of the merge, ancestry alone never justifies reporting shipped, and on a mismatch verification re-runs against the merged head before completion is declared.
|
|
31
35
|
- **Evidence references are not manifest entries.** Extract obligations using the exact `[EVIDENCE:` prefix (and the legacy local `[SCREENSHOT:` form). Exclude both the canonical `[EVIDENCE-REF: <work-item-ref> | <artifact-type>: <kebab-case-name>]` and the Lisa 2.223.0 legacy alias `[EVIDENCE-REF: <tracker-ref>: <artifact-type>: <kebab-case-name>]` from artifact lookup, missing-artifact reporting, and duplicate-name checks: either belongs to another work item and cannot satisfy this item's S14 gate. A runtime-changing leaf with references but no local claiming marker must be rejected before dispatch.
|
|
32
36
|
|
|
33
37
|
## UI Evidence Checklist (when work is UI-visible)
|
|
@@ -47,8 +51,11 @@ The checklist is tracker-agnostic — the same shape works on JIRA, GitHub Issue
|
|
|
47
51
|
5. **"What this shows" section.** Tailor to ticket type:
|
|
48
52
|
- **Bug repro:** state plainly whether the bug reproduces or not, and the most likely 1–2 reasons their retest still failed (different env, native app vs. web, stuck backend row, etc.).
|
|
49
53
|
- **Feature/UX completion:** state plainly which acceptance criteria each screenshot covers, and call out any deferred or out-of-scope surface explicitly so QA/PM doesn't have to infer.
|
|
50
|
-
6. **"
|
|
51
|
-
|
|
52
|
-
|
|
54
|
+
6. **"Artifact identity" and "Not established" sections (both required).** Right after "What this shows":
|
|
55
|
+
- `## Artifact identity` — what the evidence was collected against, as **values, not placeholders**: the repository, the `head_sha` the verification observed, the `environment`, and per artifact its `sha256` digest and `captured_at`. A heading rendered with `<unknown>` in place of a SHA is not identity.
|
|
56
|
+
- `## Not established` — **never omitted, never blank.** List in plain language what this verification did *not* prove: boundaries not exercised (e.g. "the persisted order row was never queried"), environments not tested (e.g. "checked on desktop Chrome only — not mobile Safari"), and behavior consciously out of scope (e.g. "refunds were not touched or tested"). Name the specific thing, not a category. With nothing outstanding, the heading still renders a single line: `None outstanding — reviewed`. Never state a quality check as proof of behavior — "unit tests pass" belongs here as a limit, not above as evidence.
|
|
57
|
+
7. **"What would help me confirm" (bug) / "How to QA" (feature) section.** Concrete actionable retest steps with the exact selection criteria (e.g., "pick a record whose Status column shows `—`, not `Processing` or `Pending Review`").
|
|
58
|
+
8. **Explicit invitation to be corrected.** End with a line like *"If any of the steps I listed are different from what you expected / actually did, please tell me explicitly which step I got wrong."* Non-optional — small differences (which record, which device, exact tap order, expected behavior) change everything, and naming the door open short-circuits ticket bounce-loops.
|
|
59
|
+
9. **Workflow transition** is the vendor skill's job, not yours — it'll move the ticket per the configured tracker (JIRA: Reassign to reporter for bug repro / move to the configured review status when one exists, otherwise leave it in `claimed`; GitHub: direct `claimed` → configured `done` after a successful build; Linear: equivalent state). You don't transition manually.
|
|
53
60
|
|
|
54
61
|
**Why this format:** It (a) gives the reporter a frame-by-frame they can compare against, (b) avoids the JIRA image-collapse failure mode while still working everywhere else, (c) names the most plausible discrepancies up front so the loop short-circuits, (d) explicitly opens the door to being corrected so tickets don't bounce on assumed alignment. The same mechanics that resolve a stuck bug ticket also give QA an unambiguous handoff for a freshly-built feature.
|