agent-quality-skills 1.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +293 -0
- package/bin/install.js +90 -0
- package/package.json +25 -0
- package/skills/quality-gate/SKILL.md +275 -0
- package/skills/quality-gate/references/backend-checklist.md +116 -0
- package/skills/quality-gate/references/evidence-contract.md +164 -0
- package/skills/quality-gate/references/frontend-checklist.md +98 -0
- package/skills/quality-gate/references/gates-and-fixes.md +132 -0
- package/skills/quality-gate/references/go-live-checklist.md +93 -0
- package/skills/quality-gate/references/security-checklist.md +120 -0
- package/templates/QUALITY-GATE.md +87 -0
|
@@ -0,0 +1,275 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: quality-gate
|
|
3
|
+
description: Pre-ship review gate for a codebase — run it before committing, merging, opening a PR, or releasing. Runs the repo's own lint/typecheck/test/build scripts and reports their real output, audits the diff against the project's written standards, checks the diff against the spec or ticket it came from (missing / scope creep / implemented-but-wrong), applies only mechanically-safe fixes, and returns Blocking / Major / Minor findings plus a separate no-severity Needs-validation list, ending in an Approve / Approve with Comments / Request Changes / Reject verdict. Use when someone says a module, endpoint, service, component, migration, or page is done and asks "is this ready to merge / production-ready", "review this", "does it follow our standards", "does this match the spec", or wants a security review of auth, validation, or data-access code — even without the words "skill" or "standards". Do NOT use it to write features, explain how code works, or scaffold specs.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Quality Gate
|
|
7
|
+
|
|
8
|
+
A review that a strict human reviewer would sign, produced the same way every time.
|
|
9
|
+
|
|
10
|
+
The bar is **"would a careful reviewer approve this for production?"**, not "does it
|
|
11
|
+
run". When in doubt the verdict is **Request Changes** — a flagged non-issue costs a
|
|
12
|
+
second look; a missed Blocking costs an incident.
|
|
13
|
+
|
|
14
|
+
## What this is not
|
|
15
|
+
|
|
16
|
+
- **Not a linter.** Linters are step 3 of 6, and they are the cheap part.
|
|
17
|
+
- **Not an author.** It reviews and applies mechanically-safe fixes. It does not
|
|
18
|
+
design features, rename things, or refactor.
|
|
19
|
+
- **Not a substitute for a security audit.** See `references/evidence-contract.md`
|
|
20
|
+
for where that line sits and what to do at it.
|
|
21
|
+
|
|
22
|
+
---
|
|
23
|
+
|
|
24
|
+
## Step 1 — Resolve the standards
|
|
25
|
+
|
|
26
|
+
Standards are the source of truth and they belong to the project, not to this skill.
|
|
27
|
+
Read the project's own docs when they exist; fall back to the bundled checklists when
|
|
28
|
+
they do not.
|
|
29
|
+
|
|
30
|
+
1. Look for project standards. Try, in order, and stop at the first that exists:
|
|
31
|
+
`docs/standards/`, `docs/requirements/standards/`, `standards/`, `.standards/`,
|
|
32
|
+
`CONTRIBUTING.md`. Search upward from the working directory — in a multi-repo
|
|
33
|
+
workspace the standards often live one level above the repo. Use glob/ls; do not
|
|
34
|
+
assume a path.
|
|
35
|
+
2. If found, skim the documents relevant to the files in the diff. **They override
|
|
36
|
+
anything bundled here when they conflict.**
|
|
37
|
+
3. If not found, use `references/` below. They are a faithful, general version of the
|
|
38
|
+
same rules and are enough to run a full review.
|
|
39
|
+
|
|
40
|
+
| bundled file | read it when the diff includes… |
|
|
41
|
+
|---|---|
|
|
42
|
+
| `references/backend-checklist.md` | server code: controllers, services, repositories, DTOs, entities, migrations, workers, queues |
|
|
43
|
+
| `references/frontend-checklist.md` | UI: components, pages, loaders, forms, client-side state, rendering of user input |
|
|
44
|
+
| `references/security-checklist.md` | auth, authorization, input handling, uploads, logging, secrets, outbound URLs — **almost always** |
|
|
45
|
+
| `references/evidence-contract.md` | **every run that is about to write a security finding.** What it takes for a claim to count |
|
|
46
|
+
| `references/gates-and-fixes.md` | **every run** — how to find and run the repo's real gates, and the exact safe-fix policy |
|
|
47
|
+
| `references/go-live-checklist.md` | only when asked for a release / production-readiness review, not a per-change one |
|
|
48
|
+
|
|
49
|
+
> **Project config.** If the project has a `QUALITY-GATE.md` (see `templates/`), read
|
|
50
|
+
> it first: it names the lanes, the protected branches, the queue file, and any
|
|
51
|
+
> project-specific traps. It is optional; the skill works without one.
|
|
52
|
+
|
|
53
|
+
---
|
|
54
|
+
|
|
55
|
+
## Step 2 — Determine scope
|
|
56
|
+
|
|
57
|
+
- If the user named a file, folder, or module, review that.
|
|
58
|
+
- Else, in a git repo, review the **diff**: uncommitted changes, plus the branch
|
|
59
|
+
against its base (`main`/`master`/`develop`, whichever exists).
|
|
60
|
+
- If the tree is clean and there is no branch delta, review the last commit.
|
|
61
|
+
- If there is no git and no path, ask which file or folder to review.
|
|
62
|
+
|
|
63
|
+
Review what changed **plus what it affects** — a changed DTO touches its controller,
|
|
64
|
+
its service, and any published API contract. Skip generated and vendored files.
|
|
65
|
+
|
|
66
|
+
**State the scope in one line at the top of the report**, so the reader knows exactly
|
|
67
|
+
what was and was not covered.
|
|
68
|
+
|
|
69
|
+
### Pin what the change was SUPPOSED to be
|
|
70
|
+
|
|
71
|
+
Everything else asks *is this code good?* This asks *is it the thing that was asked
|
|
72
|
+
for?* — and good code can fail it. Find the originating spec, in this order:
|
|
73
|
+
|
|
74
|
+
1. A path or spec the user named.
|
|
75
|
+
2. The project's spec directory, if it has one (`docs/specs/`, `docs/superpowers/specs/`, `spec/`).
|
|
76
|
+
3. An issue referenced in the commit messages (`#123`, `Closes #45`).
|
|
77
|
+
4. The open PR body.
|
|
78
|
+
5. The project's backlog or queue file, if it has one — an item there is a spec of sorts.
|
|
79
|
+
|
|
80
|
+
If none exists, write **"no spec available"** in the report and skip the Spec axis.
|
|
81
|
+
Do not invent one from the diff and then check the diff against it: that always
|
|
82
|
+
passes and proves nothing.
|
|
83
|
+
|
|
84
|
+
> This axis exists because the author and the reviewer are usually the same context,
|
|
85
|
+
> and a spec written an hour earlier stops being read. The characteristic failure is
|
|
86
|
+
> code that looks entirely correct on its own and is only wrong **next to the spec**.
|
|
87
|
+
|
|
88
|
+
---
|
|
89
|
+
|
|
90
|
+
## Step 3 — Run the repo's real gates
|
|
91
|
+
|
|
92
|
+
Machines catch mechanical problems faster than reading does. Run them first so your
|
|
93
|
+
reading time goes to judgment.
|
|
94
|
+
|
|
95
|
+
1. **Read the manifest and run the scripts that exist** — `package.json`,
|
|
96
|
+
`Makefile`, `pyproject.toml`, `Cargo.toml`, whatever the project uses. Typically
|
|
97
|
+
some of: lint, typecheck, test, build.
|
|
98
|
+
2. **Capture the actual output.** Report real pass/fail counts and the first real
|
|
99
|
+
errors. Never a guess.
|
|
100
|
+
3. If a script does not exist, say so (`no "lint" script found`) and move on. Absence
|
|
101
|
+
is not failure, and it is not a pass either.
|
|
102
|
+
|
|
103
|
+
Three traps that have each produced a false green, all covered in
|
|
104
|
+
`references/gates-and-fixes.md`:
|
|
105
|
+
|
|
106
|
+
- **Never pipe a gate and read `$?`** — `cmd | tail` returns `tail`'s status, not the
|
|
107
|
+
gate's. Redirect to a file and read the tool's own exit code.
|
|
108
|
+
- **Run the gates that exist, not the ones you remember.** List the manifest's scripts
|
|
109
|
+
and run each that is a gate. A suite you never ran is not a suite that passed.
|
|
110
|
+
- **An npm script is not always what CI runs.** If CI invokes a binary with flags the
|
|
111
|
+
script omits, the script can exit 0 while CI fails. Compare them.
|
|
112
|
+
|
|
113
|
+
Record each result for the Gates table.
|
|
114
|
+
|
|
115
|
+
---
|
|
116
|
+
|
|
117
|
+
## Step 4 — Audit against the standards
|
|
118
|
+
|
|
119
|
+
Read each file in scope, route it to the matching checklist, and evaluate along these
|
|
120
|
+
dimensions:
|
|
121
|
+
|
|
122
|
+
- **Correctness** — requirements met, edge cases, null/undefined, failure paths.
|
|
123
|
+
- **Security** — see the evidence contract below before writing anything here.
|
|
124
|
+
- **Architecture** — layer boundaries respected; thin controllers; logic in services;
|
|
125
|
+
persistence-only repositories; external calls where the project puts them.
|
|
126
|
+
- **API contract** — naming, versioning, consistent envelope, pagination on
|
|
127
|
+
collections, documented; no breaking change without a version.
|
|
128
|
+
- **Data** — keys, audit fields, soft delete, naming, foreign keys and indexes,
|
|
129
|
+
reversible migrations, no auto-sync in production.
|
|
130
|
+
- **Maintainability** — clear names, small functions, shallow nesting, no magic
|
|
131
|
+
values, no dead or clever code, strict types, no escape hatches.
|
|
132
|
+
- **Testing** — changed logic comes with tests; critical flows covered; no
|
|
133
|
+
sleep-based or order-dependent tests.
|
|
134
|
+
- **Performance, caching and limits** — ask the four **together** for every endpoint
|
|
135
|
+
the diff touches, not only the one the ticket named: **bounded** · **cached where
|
|
136
|
+
the project's policy says so** · **authorized on every door that reaches the same
|
|
137
|
+
write** · **rate limited where it is public**. "It was already unbounded" is not a
|
|
138
|
+
defence once the diff touches that line.
|
|
139
|
+
- **Spec fidelity** — only when Step 2 found a spec. Quote the spec line for each
|
|
140
|
+
finding: **Missing** (asked for, absent or half-done) · **Scope creep** (behaviour
|
|
141
|
+
nobody asked for — extra surface others may build on) · **Implemented but wrong**
|
|
142
|
+
(present, not what the spec described — the one that hides, because the code reads
|
|
143
|
+
fine on its own). A deviation is not automatically Blocking: judge it on
|
|
144
|
+
consequence, and if the **spec** turned out wrong, say that instead.
|
|
145
|
+
|
|
146
|
+
### Before you write a security finding
|
|
147
|
+
|
|
148
|
+
**Read `references/evidence-contract.md`.** It is short, and it decides whether what
|
|
149
|
+
you are about to write is a finding at all. Three rules:
|
|
150
|
+
|
|
151
|
+
1. **Name a boundary AND a result.** Six rows: lower-trust principal · accepted input
|
|
152
|
+
or action · intended control · crossed boundary · affected principal or resource ·
|
|
153
|
+
observed or owner-observable result. A missing best practice with no principal and
|
|
154
|
+
no crossed boundary is a hardening note, not a security finding.
|
|
155
|
+
⚠️ **The six rows are the bar for CONFIRMING. To REJECT, one line is enough — name
|
|
156
|
+
the lower-trust principal and what they gain; if you cannot, stop there.** Spending
|
|
157
|
+
the whole table to say something is not a finding is ceremony.
|
|
158
|
+
2. **`needs-validation` is a state, and it carries no severity.** When the decisive
|
|
159
|
+
fact is outside the repository — a proxy header, an IAM policy, a provider default,
|
|
160
|
+
live configuration — say so with the exact missing fact and a safe way to check it.
|
|
161
|
+
Do not guess presence *or* absence.
|
|
162
|
+
3. **Severity cannot exceed demonstrated impact.** Use the anchors in the contract.
|
|
163
|
+
If you cannot state the concrete damage, it is lower than it feels.
|
|
164
|
+
|
|
165
|
+
### Classify every finding
|
|
166
|
+
|
|
167
|
+
- **🔴 Blocking** — must fix before merge. Confirmed security gap, missing
|
|
168
|
+
authorization, data-corruption risk, undocumented breaking change, critical bug,
|
|
169
|
+
architecture violation.
|
|
170
|
+
- **🟠 Major** — should fix before release. Missing tests for changed logic, missing
|
|
171
|
+
validation, N+1 or unbounded query, incomplete error handling.
|
|
172
|
+
- **🟡 Minor** — optional. Naming, readability, small refactors.
|
|
173
|
+
- **🔍 Needs validation** — no severity. A source-grounded question blocked on a fact
|
|
174
|
+
outside the repo.
|
|
175
|
+
|
|
176
|
+
If unsure between Blocking and Major, say so in the finding and lean stricter.
|
|
177
|
+
|
|
178
|
+
---
|
|
179
|
+
|
|
180
|
+
## Step 5 — Apply safe fixes only
|
|
181
|
+
|
|
182
|
+
A fix is safe only if a formatter or linter would make it and it **cannot change
|
|
183
|
+
runtime behaviour or a public contract**.
|
|
184
|
+
|
|
185
|
+
**Apply, then list:** formatter output; linter autofix for import order, unused
|
|
186
|
+
imports, const-over-let, quote and semicolon style, trailing commas.
|
|
187
|
+
|
|
188
|
+
**Report, never apply:** anything touching logic or control flow; replacing loose
|
|
189
|
+
types with real ones; adding or changing validation; renaming public symbols,
|
|
190
|
+
endpoints, fields or enum values; security fixes; architecture moves; schema and
|
|
191
|
+
migration changes.
|
|
192
|
+
|
|
193
|
+
Re-run the affected gate after fixing so the report reflects the fixed state. Full
|
|
194
|
+
policy: `references/gates-and-fixes.md`.
|
|
195
|
+
|
|
196
|
+
---
|
|
197
|
+
|
|
198
|
+
## Step 6 — Emit the verdict
|
|
199
|
+
|
|
200
|
+
Always produce exactly this structure:
|
|
201
|
+
|
|
202
|
+
```
|
|
203
|
+
# Quality Gate — <scope, e.g. "feature/x vs develop · 6 files">
|
|
204
|
+
|
|
205
|
+
## Summary
|
|
206
|
+
<1–3 sentences: what was reviewed and overall health>
|
|
207
|
+
|
|
208
|
+
## Gates
|
|
209
|
+
| Gate | Result |
|
|
210
|
+
|-----------|---------------------------------------------------|
|
|
211
|
+
| Lint | ✅ pass (3 auto-fixed) / ❌ 5 errors / ⏭ no script |
|
|
212
|
+
| Typecheck | ✅ pass / ❌ 2 errors |
|
|
213
|
+
| Tests | ✅ 42/42 / ❌ 2 failed / ⏭ none |
|
|
214
|
+
| Build | ✅ pass / ❌ fail / ⏭ skipped |
|
|
215
|
+
|
|
216
|
+
## Spec fidelity — <spec path or issue, or "no spec available">
|
|
217
|
+
- **Missing** — <what the spec asked for and the diff does not do>, spec: "<quoted line>"
|
|
218
|
+
- **Scope creep** — <behaviour nobody asked for>
|
|
219
|
+
- **Implemented but wrong** — <present but not as described>, spec: "<quoted line>"
|
|
220
|
+
<or> Matches the spec. <or> No spec available — axis skipped.
|
|
221
|
+
|
|
222
|
+
## 🔴 Blocking (must fix before merge)
|
|
223
|
+
- `path/file.ts:24` — <defect> — violates <standard>. <smallest effective fix>.
|
|
224
|
+
<or> None.
|
|
225
|
+
|
|
226
|
+
## 🟠 Major (should fix before release)
|
|
227
|
+
<or> None.
|
|
228
|
+
|
|
229
|
+
## 🟡 Minor (optional)
|
|
230
|
+
<or> None.
|
|
231
|
+
|
|
232
|
+
## 🔍 Needs validation (no severity — the decisive fact is outside the repo)
|
|
233
|
+
- `path/file.ts:NN` — <source-grounded hypothesis>, blocked on <exact missing fact>.
|
|
234
|
+
Check by: <safe owner-observable or local step>.
|
|
235
|
+
<omit this section when empty — EXCEPT on a diff touching infrastructure, configuration
|
|
236
|
+
or deployment, where it is mandatory even as "None.". A heading that is always empty
|
|
237
|
+
trains readers to skip it.>
|
|
238
|
+
|
|
239
|
+
## Auto-fixes applied
|
|
240
|
+
<or> None.
|
|
241
|
+
|
|
242
|
+
## Verdict
|
|
243
|
+
**Request Changes** — <short justification tied to the findings above>
|
|
244
|
+
```
|
|
245
|
+
|
|
246
|
+
The verdict is one of **Approve**, **Approve with Comments**, **Request Changes**,
|
|
247
|
+
**Reject**.
|
|
248
|
+
|
|
249
|
+
Do **not** Approve if any of these hold:
|
|
250
|
+
|
|
251
|
+
- A Blocking issue exists.
|
|
252
|
+
- A **confirmed** security concern exists. A `Needs validation` entry does not block
|
|
253
|
+
on its own — it is a question, not a defect. Say plainly in the verdict that it is
|
|
254
|
+
open, and never promote it to Blocking "to be safe": doing that is what makes a
|
|
255
|
+
security list unreadable, because nobody can then tell measured from suspected.
|
|
256
|
+
- A breaking API or schema change is present without versioning and documentation.
|
|
257
|
+
- A critical test is failing.
|
|
258
|
+
- An architecture standard is violated.
|
|
259
|
+
- The diff departs from its spec and the departure is neither called out nor agreed.
|
|
260
|
+
A silent departure is not a style question: whoever reads the spec next will believe
|
|
261
|
+
something that is not true of the code.
|
|
262
|
+
|
|
263
|
+
Reserve **Reject** for a fundamentally wrong approach.
|
|
264
|
+
|
|
265
|
+
---
|
|
266
|
+
|
|
267
|
+
## Notes
|
|
268
|
+
|
|
269
|
+
- Be specific and kind. Findings are about the code's production-readiness, not the
|
|
270
|
+
author.
|
|
271
|
+
- Do not pad. "None." in a tier is a good outcome — say it plainly.
|
|
272
|
+
- **Never claim a gate passed without running it, and never invent a finding to look
|
|
273
|
+
thorough.** Every finding cites a file:line and the standard it violates.
|
|
274
|
+
- **Review is a separate pass from writing.** If you wrote the code, you are not the
|
|
275
|
+
one to approve it — and that includes the fix for a review finding.
|
|
@@ -0,0 +1,116 @@
|
|
|
1
|
+
# Backend checklist
|
|
2
|
+
|
|
3
|
+
Server-side code: controllers/handlers, services, repositories, DTOs and schemas,
|
|
4
|
+
entities, migrations, workers and queues. Stack-neutral — the project's own standards
|
|
5
|
+
win where they differ.
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## Layering
|
|
10
|
+
|
|
11
|
+
- **Handlers stay thin**: parse, authorize, delegate, shape the response. No business
|
|
12
|
+
rules, no data access, no outbound integration calls in the request path.
|
|
13
|
+
- **Services own the rules**, and own the transaction boundary.
|
|
14
|
+
- **Repositories persist**, and do nothing else.
|
|
15
|
+
- Outbound integrations go where the project puts them — a queue, a worker, an outbox.
|
|
16
|
+
A synchronous third-party call inside a request is an availability coupling; if the
|
|
17
|
+
project has a pattern for this, not using it is an architecture violation.
|
|
18
|
+
|
|
19
|
+
## Routing
|
|
20
|
+
|
|
21
|
+
- **A literal path loses to a parameterised one unless it is registered first.**
|
|
22
|
+
`/things/active` behind `/things/:id` is silently handled by `:id` — and the symptom
|
|
23
|
+
is a type error deep in the handler, not a 404. Declare literals above parameters,
|
|
24
|
+
and register a literal-path controller before the one owning the parent path.
|
|
25
|
+
- After adding a sub-path under an existing `:id` route, **call it once**. A route with
|
|
26
|
+
no caller is not a route that works.
|
|
27
|
+
|
|
28
|
+
## API contract
|
|
29
|
+
|
|
30
|
+
- Consistent naming and versioning; one success and one error envelope across the API.
|
|
31
|
+
- **Collections are paginated**, and the response reports the true total so nothing is
|
|
32
|
+
truncated silently.
|
|
33
|
+
- Every response shape documented. If the contract is published, a consumer builds from
|
|
34
|
+
the document, not from the behaviour — **a documented field that is not sent, or a
|
|
35
|
+
sent field that is not documented, is a defect in the contract**, and the cost lands
|
|
36
|
+
on someone else's week.
|
|
37
|
+
- No breaking change without a version: removing a field, making a documented field
|
|
38
|
+
optional or nullable, narrowing a type, or removing an undeprecated path.
|
|
39
|
+
- **Additive is the escape hatch.** A new field or a new endpoint is almost always
|
|
40
|
+
allowed where a change to an existing one is not.
|
|
41
|
+
|
|
42
|
+
## Data and migrations
|
|
43
|
+
|
|
44
|
+
- Stable primary keys; audit columns; soft delete where the project uses it; consistent
|
|
45
|
+
naming; foreign keys with the indexes to match.
|
|
46
|
+
- Migrations are reversible and ordered. No schema auto-sync outside development.
|
|
47
|
+
- **Moving a column moves its indexes.** A unique index defined on the old table
|
|
48
|
+
disappears quietly when the column is recreated elsewhere — check the index list, not
|
|
49
|
+
the column list.
|
|
50
|
+
- **Appending to a column with a CHECK constraint** can make an existing row
|
|
51
|
+
unrepresentable. Read the constraint before widening a value.
|
|
52
|
+
- A uniqueness rule that depends on the current time cannot live in a partial index.
|
|
53
|
+
- **`INSERT … WHERE NOT EXISTS` is not a lock.** Under read-committed two concurrent
|
|
54
|
+
transactions both pass it. Use an advisory lock or a real constraint.
|
|
55
|
+
- An upsert that emits an unconditional update will clobber a concurrent writer — make
|
|
56
|
+
the update conditional when that matters.
|
|
57
|
+
|
|
58
|
+
## Money, time and identity
|
|
59
|
+
|
|
60
|
+
- **Money is an integer in the smallest unit**, or a decimal type — never a float.
|
|
61
|
+
Sum it in the database, narrow it only at the boundary.
|
|
62
|
+
- **Timezone is a decision, not a default.** If the product's day boundary is not UTC,
|
|
63
|
+
extracting a date part from a timestamp in the server's session timezone is wrong.
|
|
64
|
+
Put the rule in one named helper.
|
|
65
|
+
- Identifiers that a person will be shown are never reissued after a race is lost.
|
|
66
|
+
|
|
67
|
+
## Performance, caching and limits
|
|
68
|
+
|
|
69
|
+
Ask the four **together** for every endpoint the diff touches, not only the one the
|
|
70
|
+
ticket named:
|
|
71
|
+
|
|
72
|
+
| | |
|
|
73
|
+
|---|---|
|
|
74
|
+
| **fast** | bounded query · no N+1 · indexed predicate · no `SELECT *` |
|
|
75
|
+
| **cached** | apply the project's policy for this data class rather than inventing one; bust on our own write; include any database or tenant discriminator **in the cache key** |
|
|
76
|
+
| **authorized** | the rule holds on every door reaching the same write |
|
|
77
|
+
| **limited** | public routes have a rate limit; outbound calls respect the partner's |
|
|
78
|
+
|
|
79
|
+
- **A consumer asking for the convenient shape is not a reason to drop a bound.** Give
|
|
80
|
+
them what their screen needs *and* a ceiling, report the true total so nothing is cut
|
|
81
|
+
silently, and say what the ceiling is.
|
|
82
|
+
- No per-row lookup inside a loop on a paginated endpoint.
|
|
83
|
+
- Check the query count on any list or dashboard route — those are the hot ones.
|
|
84
|
+
|
|
85
|
+
## Caching hazards
|
|
86
|
+
|
|
87
|
+
- **One cache instance shared across environments needs the environment in the key**,
|
|
88
|
+
or a value from a dropped database is served to a live one.
|
|
89
|
+
- **Bump a version segment in the key when a value's shape changes**, or a stale entry
|
|
90
|
+
deserialises into a shape the code no longer handles.
|
|
91
|
+
- A cache-aside helper with single-flight must not be re-entered on its own key by its
|
|
92
|
+
own factory: it hands the factory the promise it is producing and the request never
|
|
93
|
+
settles.
|
|
94
|
+
|
|
95
|
+
## Errors
|
|
96
|
+
|
|
97
|
+
- Fail closed. A degraded dependency should make the answer slower or narrower, never
|
|
98
|
+
more permissive.
|
|
99
|
+
- Distinguish "the store is unavailable" (often survivable) from "the rule says no"
|
|
100
|
+
(never survivable).
|
|
101
|
+
|
|
102
|
+
## Testing
|
|
103
|
+
|
|
104
|
+
- Changed business logic ships with a test. Missing is Major — Blocking for auth,
|
|
105
|
+
money, permissions, and anything on a published contract.
|
|
106
|
+
- **Unit specs are where a new dependency surfaces.** A hand-built collaborator in a
|
|
107
|
+
unit spec is the only written record of that class's dependency shape; an
|
|
108
|
+
integration test wires the real thing and stays green.
|
|
109
|
+
- **Shared-database suites must seed what they assert on**, and purge it. A suite that
|
|
110
|
+
reads another suite's leftovers passes locally and fails on a fresh database.
|
|
111
|
+
- **An assertion that loops over a list passes on an empty list.** Put a length
|
|
112
|
+
assertion ahead of every such loop — without it the test cannot fail, which is the
|
|
113
|
+
one thing a security test must be able to do.
|
|
114
|
+
- Fixtures must be unique **by construction**, not by luck. A derived random value with
|
|
115
|
+
a small space collides eventually, and the failure surfaces inside a helper where it
|
|
116
|
+
reads as a bug in the code under test.
|
|
@@ -0,0 +1,164 @@
|
|
|
1
|
+
# Evidence contract for security findings
|
|
2
|
+
|
|
3
|
+
A checklist answers *does the code do X?* This file answers the other half: **what it
|
|
4
|
+
takes for a claim about X to count.**
|
|
5
|
+
|
|
6
|
+
It exists because the expensive failure in review is rarely a missed check. It is a
|
|
7
|
+
claim that says more than it measured — a suspicion written in the same ink as a
|
|
8
|
+
proven defect, so that a week later nobody can tell them apart, and the real one is
|
|
9
|
+
buried in the pile.
|
|
10
|
+
|
|
11
|
+
The three rules are adapted from Cloudflare's `security-audit` skill
|
|
12
|
+
(MIT, <https://github.com/cloudflare/security-audit-skill>), which is the companion to
|
|
13
|
+
this gate. See "Where this gate stops" at the bottom.
|
|
14
|
+
|
|
15
|
+
---
|
|
16
|
+
|
|
17
|
+
## Rule 1 — A finding needs a boundary AND a result
|
|
18
|
+
|
|
19
|
+
Fill in all six, or it is not a security finding:
|
|
20
|
+
|
|
21
|
+
| | |
|
|
22
|
+
|---|---|
|
|
23
|
+
| **lower-trust principal** | who is doing this — an anonymous caller, a basic user, a suspended account, another tenant's admin |
|
|
24
|
+
| **accepted input or action** | what they send or do |
|
|
25
|
+
| **intended control** | the guard or rule that is supposed to stop it |
|
|
26
|
+
| **crossed boundary** | which line that control was drawing |
|
|
27
|
+
| **affected principal or resource** | whose data or capability moves |
|
|
28
|
+
| **observed or owner-observable result** | what actually happens — not what could |
|
|
29
|
+
|
|
30
|
+
### The six rows are the bar for CONFIRMING — rejecting takes one line
|
|
31
|
+
|
|
32
|
+
**To reject a candidate, ask one question: name the lower-trust principal and
|
|
33
|
+
what they gain.** If you cannot, it is a hardening note. Stop there.
|
|
34
|
+
|
|
35
|
+
The table is what a **confirmed** finding must carry, and that is where it earns
|
|
36
|
+
its cost — somebody is going to act on the answer. Filling six rows in order to
|
|
37
|
+
throw a candidate away is heavier than the throwing-away deserves, and a rule
|
|
38
|
+
that costs more than it returns is one people quietly stop using.
|
|
39
|
+
|
|
40
|
+
### What this excludes, on purpose
|
|
41
|
+
|
|
42
|
+
- **A missing best practice with no principal and no crossed boundary.** "Tokens are
|
|
43
|
+
kept in `localStorage`" names nobody and crosses nothing on its own. It becomes a
|
|
44
|
+
finding when you can say who reads them and what they then reach. Until then it is
|
|
45
|
+
a hardening note, and filing it as a security finding makes the real ones harder to
|
|
46
|
+
see.
|
|
47
|
+
- **Guessed deployment behaviour.** A proxy, provider default, browser policy or IAM
|
|
48
|
+
rule that is not in the repository is not evidence in either direction. That is what
|
|
49
|
+
Rule 2 is for.
|
|
50
|
+
- **A generic crash with no security outcome.**
|
|
51
|
+
- **Self-impact.** A caller degrading only their own session has not crossed a
|
|
52
|
+
boundary.
|
|
53
|
+
|
|
54
|
+
---
|
|
55
|
+
|
|
56
|
+
## Rule 2 — `needs-validation` is a STATE, and it carries no severity
|
|
57
|
+
|
|
58
|
+
Three states, and keeping them apart is the whole point:
|
|
59
|
+
|
|
60
|
+
| state | means |
|
|
61
|
+
|---|---|
|
|
62
|
+
| **confirmed** | the six rows above are filled in from source, or from a bounded local run |
|
|
63
|
+
| **needs-validation** | a specific, source-grounded hypothesis whose **decisive fact is outside** what the repository can show |
|
|
64
|
+
| **rejected** | traced, and the boundary holds |
|
|
65
|
+
|
|
66
|
+
🔴 **`needs-validation` is not "a confirmed finding I am less sure about".** It gets
|
|
67
|
+
**no severity at all** — not a lower one. A severity on an unverified item is how a
|
|
68
|
+
list of *things to go and check* silently becomes a list of *holes*.
|
|
69
|
+
|
|
70
|
+
A `needs-validation` entry must carry:
|
|
71
|
+
|
|
72
|
+
1. the exact missing fact (`is the load balancer setting X-Forwarded-For, or is the
|
|
73
|
+
app trusting a client-supplied one?`), and
|
|
74
|
+
2. a safe way to establish it — an owner-observable check or a bounded local test.
|
|
75
|
+
Never a probe against production or shared infrastructure.
|
|
76
|
+
|
|
77
|
+
**Write it down even when you cannot resolve it.** An unrecorded question comes back
|
|
78
|
+
as an incident; a recorded one comes back as a task.
|
|
79
|
+
|
|
80
|
+
⚠️ **Print the `needs-validation` section only when it has something in it** —
|
|
81
|
+
except on a diff touching infrastructure, configuration or deployment, where it
|
|
82
|
+
is mandatory even as "None.", because that is where the outside-the-repo facts
|
|
83
|
+
live and "None." is then a real statement. Elsewhere a standing empty heading
|
|
84
|
+
trains readers to skip the section, which defeats the rule it exists to serve.
|
|
85
|
+
|
|
86
|
+
---
|
|
87
|
+
|
|
88
|
+
## Rule 3 — Severity cannot exceed demonstrated impact
|
|
89
|
+
|
|
90
|
+
| | |
|
|
91
|
+
|---|---|
|
|
92
|
+
| **critical** | an unauthenticated actor gains code execution, full data-store access, or takeover of arbitrary accounts |
|
|
93
|
+
| **high** | an explicit control is **fully defeated** with real consequence — auth bypass, cross-tenant read or write, stored script execution affecting other users, authenticated code execution, unauthenticated remote stop of a shared service |
|
|
94
|
+
| **medium** | a real boundary violation with limited blast radius or uncommon preconditions |
|
|
95
|
+
| **low** | non-secret internals disclosed, or sustained effort for minimal gain |
|
|
96
|
+
| **informational** | confirmed but minimal — useful mainly as a step inside a larger finding |
|
|
97
|
+
|
|
98
|
+
**The high/medium line:** does the demonstrated result *fully defeat* the control, or
|
|
99
|
+
only weaken it?
|
|
100
|
+
|
|
101
|
+
**If you cannot state the concrete damage, the severity is lower than it feels.**
|
|
102
|
+
|
|
103
|
+
> This governs how a finding is **described**. It does not soften the merge rule: a
|
|
104
|
+
> confirmed security gap is Blocking wherever it lands on this table.
|
|
105
|
+
|
|
106
|
+
---
|
|
107
|
+
|
|
108
|
+
## Rule 4 — Recommend the smallest effective fix
|
|
109
|
+
|
|
110
|
+
Name the invariant the code must enforce, and the narrowest change that enforces it
|
|
111
|
+
**at the last trusted decision point**. A repository-relative edit plus a regression
|
|
112
|
+
test, not generic hardening advice.
|
|
113
|
+
|
|
114
|
+
If the same rule is enforced at several doors, say so — a fix at one door while the
|
|
115
|
+
others still reach the same write is a fix that moves the defect rather than closing
|
|
116
|
+
it.
|
|
117
|
+
|
|
118
|
+
---
|
|
119
|
+
|
|
120
|
+
## Where this gate stops, and the audit begins
|
|
121
|
+
|
|
122
|
+
This gate reviews **a diff**. A security audit surveys **a system**, and they are not
|
|
123
|
+
the same activity or the same cost.
|
|
124
|
+
|
|
125
|
+
Use the companion skill's **guidance mode** from here: when the diff calls for it,
|
|
126
|
+
read the ONE domain file that matches. Reading a reference is not running an audit.
|
|
127
|
+
|
|
128
|
+
| the diff touches | read |
|
|
129
|
+
|---|---|
|
|
130
|
+
| auth, sessions, tokens, CORS, headers, route guards | `WEB-PROTOCOL-AND-AUTH.md` |
|
|
131
|
+
| tenancy, soft delete, exports, retention, cross-account reads | `DATA-ISOLATION-AND-LIFECYCLE.md` |
|
|
132
|
+
| cloud runtime, IAM, buckets, caches, manifests, environment | `CLOUD-AND-DEPLOYMENT.md` |
|
|
133
|
+
| rendering user input, uploads, CSP, client-side state | `CLIENT-SIDE.md` |
|
|
134
|
+
| pagination, unbounded queries, rate limits, timeouts | `RESOURCE-EXHAUSTION-AND-AVAILABILITY.md` |
|
|
135
|
+
| dependencies, build, release, signing | `SUPPLY-CHAIN-AND-RELEASE.md` |
|
|
136
|
+
|
|
137
|
+
🔴 **Do NOT run that skill's six-phase workflow from inside this gate.** Its unit of
|
|
138
|
+
cost is an *agent invocation*: even its `quick` profile is four reconnaissance calls
|
|
139
|
+
plus a hunter wave, a critic, and one or two verifiers per candidate. A per-change
|
|
140
|
+
gate cannot afford that, and the skill says so itself — loading it does not authorize
|
|
141
|
+
the workflow.
|
|
142
|
+
|
|
143
|
+
**Schedule full audit mode as its own activity**, ideally before an external
|
|
144
|
+
penetration test so the cheap findings are already fixed.
|
|
145
|
+
|
|
146
|
+
⚠️ **Know this before the run, not during it:** that skill requires an OS-enforced
|
|
147
|
+
sandbox — no network, empty environment, hard CPU/memory/wall-clock limits — before
|
|
148
|
+
executing any target code, and states that without every control you do not execute
|
|
149
|
+
it. Most machines do not have one. A run without it is **source-only**, and every
|
|
150
|
+
dynamic claim comes back `needs-validation`. That is the honest result, and it is
|
|
151
|
+
exactly why Rule 2 has to be in place *before* the audit rather than after it.
|
|
152
|
+
|
|
153
|
+
---
|
|
154
|
+
|
|
155
|
+
## Anti-patterns
|
|
156
|
+
|
|
157
|
+
1. A checklist deviation presented as a vulnerability.
|
|
158
|
+
2. Defence-in-depth advice with no reachable boundary violation.
|
|
159
|
+
3. Probing live or shared environments where bounded local evidence would do.
|
|
160
|
+
4. Guessing provider, proxy, browser, identity or deployment behaviour absent from source.
|
|
161
|
+
5. Treating same-principal authority or self-impact as a crossed boundary.
|
|
162
|
+
6. Reporting an effect stronger than the one observed.
|
|
163
|
+
7. Assigning a severity to a `needs-validation` record.
|
|
164
|
+
8. Letting the prose and the structured findings disagree.
|
|
@@ -0,0 +1,98 @@
|
|
|
1
|
+
# Frontend checklist
|
|
2
|
+
|
|
3
|
+
UI code: components, pages, route loaders, forms, client state, and anywhere user
|
|
4
|
+
input is rendered. Stack-neutral — the project's own standards win where they differ.
|
|
5
|
+
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
## The frontend is never authoritative
|
|
9
|
+
|
|
10
|
+
- Every rule the UI expresses is also enforced on the server. A hidden button is UX,
|
|
11
|
+
not a permission.
|
|
12
|
+
- **A control that reports success and changes nothing is Blocking**, not Minor — even
|
|
13
|
+
when the broken behaviour predates the change. Adding the control is what makes it
|
|
14
|
+
reachable, so the change owns it.
|
|
15
|
+
- Where the server refuses, show **the server's own reason**. A message invented in the
|
|
16
|
+
client drifts from the rule the server actually enforces, and then two places disagree
|
|
17
|
+
about what happened.
|
|
18
|
+
|
|
19
|
+
## Write mappers — the blind spot
|
|
20
|
+
|
|
21
|
+
The single highest-value area in a UI review, because a mocked or faked API layer
|
|
22
|
+
**cannot see it**: a mock never receives a request body, so a mapper that *drops* a
|
|
23
|
+
field and one that *clears* it look identical from the mock.
|
|
24
|
+
|
|
25
|
+
- **`undefined` is not "empty" — it is "don't touch this".** Serialisation drops the
|
|
26
|
+
key, so a partial update leaves whatever the record had. If the intent is *clear this
|
|
27
|
+
field*, send the explicit empty value the API accepts.
|
|
28
|
+
- **Never send an empty string to mean empty** where the API's empty is null: `""` is
|
|
29
|
+
accepted and stored, then reaches consumers as a string rather than an absence, and
|
|
30
|
+
downstream renderers throw on it.
|
|
31
|
+
- **Absent and empty-collection are different requests.** If the API treats a missing
|
|
32
|
+
key as "leave alone" and an empty array as "remove all", the mapper must be able to
|
|
33
|
+
express both, and the difference must be a distinct code path — not a falsy check.
|
|
34
|
+
- **Anything mapping UI state onto a request body needs a test that stubs the transport
|
|
35
|
+
and asserts on the serialised body.** Not on the mock. That test is the only thing
|
|
36
|
+
that can distinguish the two failures above.
|
|
37
|
+
|
|
38
|
+
## Read mappers
|
|
39
|
+
|
|
40
|
+
- Map at the boundary; do not pass a raw response through typed as the domain object.
|
|
41
|
+
A cast type-checks and validates nothing, and the field you add next arrives unmapped.
|
|
42
|
+
- Decide what a malformed value defaults to, and pick the default that **fails toward
|
|
43
|
+
the recoverable side** — an over-permissive screen the server then refuses beats a
|
|
44
|
+
screen that cannot be used.
|
|
45
|
+
|
|
46
|
+
## Loading, empty and failure are three states, not two
|
|
47
|
+
|
|
48
|
+
- **A failed read is not an empty result.** Many data libraries report "not loading"
|
|
49
|
+
once a request errors, so keying only on a loading flag draws "there is nothing here"
|
|
50
|
+
over data that exists and the user cannot tell.
|
|
51
|
+
- A paused query — offline, or a network-mode guard — is a fourth case that looks like
|
|
52
|
+
neither. Use an explicit "has anything arrived?" check rather than inferring it.
|
|
53
|
+
- Put the error branch **first**, so a 403 reads as "you may not see these" rather than
|
|
54
|
+
"there are none".
|
|
55
|
+
|
|
56
|
+
## Lists and pagination
|
|
57
|
+
|
|
58
|
+
- **One request with no limit is not "the whole list" — it is the first page,
|
|
59
|
+
silently.** Either page through to the server's reported total, or ask for a bound
|
|
60
|
+
and show it.
|
|
61
|
+
- If the screen prints a count, it must come from a **different source** than the array
|
|
62
|
+
it labels. A count taken from the rendered array agrees with the screen by
|
|
63
|
+
construction and can never catch a truncation.
|
|
64
|
+
|
|
65
|
+
## Rendering user input
|
|
66
|
+
|
|
67
|
+
- Escape everything that came from a person. Where a page renders both authored content
|
|
68
|
+
and submitted content, **check which path each field takes** — the usual defect is
|
|
69
|
+
one field crossing from the escaped path to the authored one, and it is stored XSS on
|
|
70
|
+
a page colleagues read.
|
|
71
|
+
|
|
72
|
+
## Forms
|
|
73
|
+
|
|
74
|
+
- Validate with the same rules the server uses, and treat the server as the decider.
|
|
75
|
+
- Trim before length rules, not after, or whitespace passes a minimum and stores empty.
|
|
76
|
+
- Surface field-level server errors next to their fields.
|
|
77
|
+
|
|
78
|
+
## Internationalisation
|
|
79
|
+
|
|
80
|
+
- No user-visible string hardcoded when the project has a message catalogue.
|
|
81
|
+
- Adding a key means adding it to **every** locale; a missing translation should be
|
|
82
|
+
visible in a check, not at runtime.
|
|
83
|
+
|
|
84
|
+
## Accessibility and layout
|
|
85
|
+
|
|
86
|
+
- Interactive controls are reachable by keyboard and carry an accessible name.
|
|
87
|
+
- Icon-only buttons need a label.
|
|
88
|
+
- **A page that loads is not a page that renders.** Every link resolving is not evidence
|
|
89
|
+
of a working screen — a layout can collapse to an unusable strip while every request
|
|
90
|
+
returns 200. When a page is reported broken, open it and measure the container.
|
|
91
|
+
|
|
92
|
+
## Testing
|
|
93
|
+
|
|
94
|
+
- Test the mapping and the state machine; do not re-test the framework.
|
|
95
|
+
- **A test of on-screen copy must anchor the claim it makes.** Asserting a string exists
|
|
96
|
+
proves the string exists, not that it is shown in the right state.
|
|
97
|
+
- Prefer a walk through the real screen for a flow bug; an assertion can agree with a
|
|
98
|
+
broken screen when both were written from the same wrong assumption.
|