agent-quality-skills 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,132 @@
1
+ # Mechanical gates and the safe-fix policy
2
+
3
+ How to find and run the repo's real gates, and exactly what may be auto-fixed. The
4
+ goal: let tooling catch mechanical problems so reading time goes to judgment, and only
5
+ ever apply a change that cannot alter behaviour or a public contract.
6
+
7
+ ---
8
+
9
+ ## Detect, don't assume
10
+
11
+ Read the project's manifest and run the scripts that **actually exist**. Do not run a
12
+ command that is not defined — report it as absent instead. Run long gates in the
13
+ background so review continues while they do.
14
+
15
+ ```bash
16
+ # node
17
+ node -e "console.log(Object.keys(require('./package.json').scripts||{}).join('\n'))"
18
+ # make
19
+ grep -E '^[a-zA-Z_-]+:' Makefile
20
+ # python / rust / go
21
+ ls pyproject.toml tox.ini Cargo.toml Makefile go.mod 2>/dev/null
22
+ ```
23
+
24
+ Then run each one that is a gate. Typical shapes:
25
+
26
+ | gate | look for |
27
+ |---|---|
28
+ | lint | `lint`, `lint:ci`, `check`, `ruff`, `clippy`, `golangci-lint` |
29
+ | format | `format`, `fmt`, `format:check` — prefer the **check** variant for reporting |
30
+ | typecheck | `typecheck`, `check`, `tsc --noEmit`, `mypy`, `svelte-check` |
31
+ | test | `test`, `test:unit`; integration and e2e separately |
32
+ | build | `build`, `compile` |
33
+
34
+ **Package manager and monorepo:** honour the lockfile (`pnpm`/`yarn`/`npm`/`bun`). In
35
+ a monorepo, scope to the changed workspace rather than building the world.
36
+
37
+ **Integration and e2e suites** often need a database, a cache, a browser, or a network
38
+ service. If one cannot run cleanly here, **say so and separate it from genuine
39
+ failures** — a false red costs as much trust as a false green.
40
+
41
+ ---
42
+
43
+ ## Three traps that each produced a false green
44
+
45
+ ### 1. Never pipe a gate and then read `$?`
46
+
47
+ ```bash
48
+ npm test | tail -5 ; echo $? # ← this is TAIL's exit code. Always 0.
49
+ ```
50
+
51
+ A failing suite reads as a pass. Redirect and read the tool's own status:
52
+
53
+ ```bash
54
+ npm test > /tmp/test.log 2>&1 ; echo "exit=$?"
55
+ ```
56
+
57
+ The same applies to a trailing `echo`, and to any wrapper that reports the status of
58
+ the last command in a chain rather than the gate's.
59
+
60
+ ### 2. Run the gates that exist, not the ones you remember
61
+
62
+ Enumerate the manifest every time. A suite you never ran is not a suite that passed,
63
+ and the one you forgot is usually the one that would have gone red — a change that
64
+ adds a dependency to a class breaks the **unit** spec that hand-builds it while
65
+ integration tests, which wire the real thing, stay green.
66
+
67
+ ### 3. The script is not always what CI runs
68
+
69
+ If CI invokes a binary with flags the script omits, the script can exit 0 while CI
70
+ fails — for example a docs or lint checker whose npm script lacks the `--check` flag
71
+ and therefore always succeeds. Compare the script body against the CI invocation.
72
+
73
+ ---
74
+
75
+ ## Reporting gate results
76
+
77
+ - Report the **real** counts from the output: `Tests: ❌ 2 failed (40 passed)`, with
78
+ the first failing names. Never guess.
79
+ - A missing script is `⏭ no "lint" script found` — not a failure, and not a pass.
80
+ - **Sanity-check the test count against the baseline.** If the diff adds four tests,
81
+ the total should move by four. A total that did not move, or moved by the wrong
82
+ amount, means a suite silently did not run — which a green exit code will not tell
83
+ you. Measure the baseline on the base branch rather than recalling it.
84
+
85
+ ---
86
+
87
+ ## Safe-fix policy
88
+
89
+ A fix is **safe to auto-apply** only if a formatter or linter would make it **and** it
90
+ cannot change runtime behaviour or a public contract.
91
+
92
+ ### Auto-fix — apply, then list each one with its file
93
+
94
+ - The project formatter (`prettier --write`, `ruff format`, `gofmt`, `cargo fmt`).
95
+ - Linter autofix of safe rules: import ordering, unused-import removal,
96
+ const-over-let, quote / semicolon / trailing-comma style, simple reflows.
97
+
98
+ ### Report only — never auto-apply
99
+
100
+ - Any change to logic, control flow, or behaviour.
101
+ - Replacing loose types with concrete ones; adding or changing validation schemas.
102
+ - Renaming public symbols, endpoints, response fields, or enum values — a consumer
103
+ depends on every one of them.
104
+ - Security changes: adding authorization, sanitisation, rate limiting.
105
+ - Architecture moves: taking logic out of a controller, introducing a worker path.
106
+ - Any schema or migration change.
107
+ - Deleting code that looks dead, unless you have confirmed it is unreferenced. A
108
+ removed export can be a breaking change; when unsure, report instead of deleting.
109
+
110
+ ### After auto-fixing
111
+
112
+ Re-run the affected gate so the report reflects the corrected state. If an autofix
113
+ touched something logic-adjacent, review that diff before keeping it.
114
+
115
+ ---
116
+
117
+ ## Proving a test is load-bearing
118
+
119
+ When the change adds a test, the report is stronger if the test is shown to be able to
120
+ fail. Neutralise the fix, confirm the test goes red **for the right reason**, restore
121
+ it. A test that stays green under a mutation of the code it claims to cover is proving
122
+ nothing, however well it reads.
123
+
124
+ Two things worth stating when you do this:
125
+
126
+ - **which** cases went red, counted — not "the test fails", but which ones and how
127
+ many; and
128
+ - **how** they failed — an assertion mismatch and a hang are different evidence, and
129
+ the one that matches the production symptom is the stronger claim.
130
+
131
+ Undo a mutation from a saved copy, never with a checkout that could take unrelated
132
+ work with it.
@@ -0,0 +1,93 @@
1
+ # Go-live / release-readiness checklist
2
+
3
+ Use this **only** when asked for a release, go-live or production-readiness review —
4
+ never for a routine per-change review.
5
+
6
+ Produce a report grouped by priority, each item ✅ / ❌ / ⏭ (n/a) with a one-line note,
7
+ then a **Go / No-Go**. Any failing P0 → **No-Go**.
8
+
9
+ > ⚠️ **Mark ✅ only from evidence you gathered in this run.** An item copied forward
10
+ > from a previous release, or taken from a status page, is ⏭ with a note saying who
11
+ > owns it — not ✅. The single most common way a release checklist misleads is a tick
12
+ > that records somebody's memory rather than a check.
13
+
14
+ ---
15
+
16
+ ## P0 — must pass
17
+
18
+ **Security** — authentication verified · authorization verified on every protected
19
+ action · no secret in source or logs · production secrets in a secret store ·
20
+ input validation · upload validation · rate limits on public routes · CORS reviewed ·
21
+ transport encrypted · credentials and tokens absent from logs.
22
+
23
+ **Data** — production backup taken **and a restore tested** · every migration reviewed
24
+ and run both directions · constraints and indexes verified · connection pool sized ·
25
+ rollback documented.
26
+
27
+ **Infrastructure** — environment and variables verified · images built and versioned ·
28
+ health checks answering · certificates valid · DNS correct · firewall reviewed.
29
+
30
+ **Observability** — service, data-store and worker monitoring on · error alerting on ·
31
+ **alert delivery tested**, not just configured.
32
+
33
+ **Logging** — structured · correlation ids present · no sensitive field logged.
34
+
35
+ **Async and integrations** — queues, retries and dead-letter paths configured · workers
36
+ connected · each external integration's credentials, connectivity, retry and failure
37
+ handling exercised.
38
+
39
+ **Application flows** — every critical user journey tested end to end, named
40
+ individually rather than as a group.
41
+
42
+ **Deployment** — release notes · deployment plan · **rollback plan** · named owner ·
43
+ business approval.
44
+
45
+ **Delivery channel** — ⚠️ the one most often missing from a plan that tracks only code
46
+ and data. Can the built artefact actually reach users? Store or distribution account in
47
+ good standing, signing identity valid, build toolchain at the version the store
48
+ requires, and somebody named as its owner. Verify it by pushing a build through the
49
+ real path, not by assuming — this is regularly discovered on the day of release,
50
+ because nothing in the development process touches it.
51
+
52
+ ---
53
+
54
+ ## P1 — should pass
55
+
56
+ **Testing** — unit, integration and end-to-end green · regression done · user
57
+ acceptance done.
58
+
59
+ **Performance** — response times measured · slow queries reviewed · N+1 checked ·
60
+ indexes confirmed · queue throughput reviewed. Measured numbers, not impressions.
61
+
62
+ **Localisation** — every locale's content verified · switching and fallback verified.
63
+
64
+ **Documentation** — API reference regenerated **and published** — the generated file in
65
+ the repository is not what consumers read · deployment docs current.
66
+
67
+ **Auditability** — audit log covers the actions that matter, including the value
68
+ before a change, and someone can read it.
69
+
70
+ **CI/CD** — build, test, deploy and rollback pipelines each exercised.
71
+
72
+ ---
73
+
74
+ ## P2 — nice to have
75
+
76
+ Caching and asset optimisation · runbook, on-call and incident docs · disaster-recovery
77
+ rehearsal with a stated objective · capacity and scaling reviewed.
78
+
79
+ ---
80
+
81
+ ## Final Go / No-Go
82
+
83
+ - [ ] All P0 complete
84
+ - [ ] Outstanding risks written down, with an owner each
85
+ - [ ] Stakeholders informed
86
+ - [ ] Rollback approved
87
+
88
+ **Recommendation: Go / No-Go**, with justification.
89
+
90
+ > 🔒 **Do not publish an open security item.** A vulnerability already fixed is good
91
+ > material for a release note — it shows the review is real. One still open tells a
92
+ > reader where to push before it is closed. Track those internally and say only that
93
+ > internal security follow-ups exist, if it must be mentioned at all.
@@ -0,0 +1,120 @@
1
+ # Security checklist (cross-cutting)
2
+
3
+ Read **`evidence-contract.md` first** if you are about to write a finding. This file
4
+ is the *what to look at*; that one is *what makes a claim count*.
5
+
6
+ Security is never deferred: a **confirmed** gap is Blocking, full stop. A suspicion is
7
+ `needs-validation` and carries no severity.
8
+
9
+ Principles: least privilege · defence in depth · secure by default · explicit
10
+ authorization · fail securely. Never trust user input, client-side validation, or an
11
+ external integration. **Every control is enforced server-side.**
12
+
13
+ ---
14
+
15
+ ## Authentication
16
+
17
+ - Short-lived access credentials; refresh rotated; signature **and** expiry verified.
18
+ - Login, logout, refresh and password-reset paths all exist and all agree.
19
+ - Passwords hashed with a memory-hard or adaptive KDF (argon2, bcrypt, scrypt) — never
20
+ plaintext, never a bare fast hash, never logged.
21
+ - A token minted for one purpose must not be accepted for another. If the system issues
22
+ scoped or single-purpose tokens, the verifier checks the scope, not just the
23
+ signature.
24
+
25
+ ## Authorization
26
+
27
+ - Enforced on the **server**, at route *and* resource level. A frontend check is UX.
28
+ - **Ask it as: who writes this value, who can change it, and who can see it?** — not
29
+ "does the check exist?". The check usually exists.
30
+ - **Every door that reaches the same write must carry the same rule.** A guard on the
31
+ endpoint in the ticket, with a sibling route reaching the same mutation unguarded, is
32
+ a gap — and it is the most common one.
33
+ - Hiding a control in the client is not hiding the data: if it is on the wire, it has
34
+ leaked.
35
+
36
+ ## Input validation and injection
37
+
38
+ - Validate body, query, path params and uploads at the boundary, with a schema.
39
+ - **SQL:** parameterised queries only. String-interpolated SQL with user input is
40
+ Blocking.
41
+ - **XSS:** escape on output; never render untrusted HTML. Where a page renders both
42
+ authored content and user-submitted content, **check which path each field takes** —
43
+ the usual defect is one field crossing from the escaped path to the authored one.
44
+ - **SSRF:** never fetch a user-supplied URL without validating scheme and host and
45
+ blocking internal ranges.
46
+ - **Deserialisation and templating:** no user input reaching an evaluator.
47
+
48
+ ## Rate limiting and abuse
49
+
50
+ - Public and unauthenticated routes are limited. Login, one-time codes and
51
+ password-reset need the tightest.
52
+ - Outbound calls respect the **partner's** limit too, not only ours.
53
+ - A counter that does not move for one class of caller is itself an oracle — check that
54
+ a refusal path and a success path are not distinguishable by timing or by a side
55
+ effect when the intent is that they should not be.
56
+
57
+ ## File upload
58
+
59
+ - Validate size, declared type **and** actual content; allow-list rather than deny-list.
60
+ - Never execute an upload; store outside any executable path; serve private files
61
+ through a signed, expiring URL rather than a public one.
62
+ - **Keep public and private storage physically separate**, and make the separation
63
+ structural — two ports or two buckets, not one flag. A flag is one wrong argument
64
+ away from publishing an identity document; a missing method is not.
65
+
66
+ ## Secrets
67
+
68
+ - No secret in source: signing keys, database and cache passwords, integration
69
+ credentials, API keys. A hardcoded secret is Blocking.
70
+ - Never commit environment files, private keys or tokens. Run a secret scanner over
71
+ history, not just the diff.
72
+ - A default value for a secret that is only refused in production is a secret in source.
73
+
74
+ ## Logging
75
+
76
+ - Structured, with a correlation id and the acting principal.
77
+ - **Never log** passwords, tokens, authorization headers, one-time codes, full card or
78
+ identity numbers. Logging any of these is Blocking.
79
+ - Log the id, not the person: a user id is fine, an email address in a log is a
80
+ retention problem.
81
+
82
+ ## Error handling
83
+
84
+ - No stack trace, SQL, internal path or secret in a response. Generic message out,
85
+ detail to the log.
86
+ - **The message must be true.** A handler that answers "invalid credentials" from a
87
+ path that never checked a credential sends the client to re-prompt for something that
88
+ was correct, and sends whoever reads the log hunting a bug that does not exist.
89
+
90
+ ## Infrastructure, when visible in the diff
91
+
92
+ - Transport encrypted; only required ports exposed.
93
+ - Data stores authenticated and not publicly reachable.
94
+ - If the app trusts a forwarded client IP, something upstream must be setting it — if
95
+ that cannot be established from the repository, it is `needs-validation`, with the
96
+ consequence stated (per-IP limits and any IP kept as evidence are wrong).
97
+
98
+ ## Dependencies
99
+
100
+ - No unused packages. Known-vulnerability check before release. A pinned-backwards
101
+ version to dodge a break can drag in advisories — check before pinning down.
102
+
103
+ ---
104
+
105
+ ## Quick triage
106
+
107
+ **Blocking** — missing authn/authz on a reachable action · secret or token in code or
108
+ logs · interpolated SQL · untrusted HTML rendered · arbitrary URL fetched · uploads
109
+ executed or exposed · responses leaking internals · **a control that reports success
110
+ and changes nothing** (adding the button is what makes it reachable, so the change owns
111
+ it, even when the broken behaviour predates it).
112
+
113
+ **Major** — missing rate limit on a sensitive public route · weak upload validation ·
114
+ missing validation not yet exploitable · missing audit trail on a sensitive action.
115
+
116
+ **Minor** — hardening suggestions, log-field completeness, naming.
117
+
118
+ **Needs validation** — anything whose decisive fact is a proxy, provider, identity or
119
+ deployment behaviour not present in the repository. No severity. State the missing fact
120
+ and a safe way to check it.
@@ -0,0 +1,87 @@
1
+ # Quality gate — project configuration
2
+
3
+ Copy this to your repository root as `QUALITY-GATE.md` and fill it in. It is
4
+ **optional**: the gate works without it, using the bundled checklists and whatever
5
+ scripts your manifest defines. This file is how a project overrides those defaults and
6
+ records the traps a reviewer cannot infer from the code.
7
+
8
+ Delete every section you do not need. An empty heading is worse than no heading — it
9
+ reads as "nothing here", which is a claim.
10
+
11
+ ---
12
+
13
+ ## Standards
14
+
15
+ Where the written standards live, if not in one of the conventional locations:
16
+
17
+ ```
18
+ docs/standards/
19
+ ```
20
+
21
+ Which ones cover what, if it is not obvious from the filenames.
22
+
23
+ ---
24
+
25
+ ## Lanes
26
+
27
+ If not every change needs the same depth of review, say so here — otherwise the gate
28
+ treats them all as full.
29
+
30
+ | lane | what falls in it | when the gate runs |
31
+ |---|---|---|
32
+ | **full** | auth · money · permissions · migrations · anything on a published contract | before merge |
33
+ | **fast** | internal screens · copy · docs · refactor with no behaviour change · test-only | batched |
34
+
35
+ > When unsure, it is **full**. A change that looks internal but alters what an external
36
+ > consumer renders is full lane.
37
+
38
+ ---
39
+
40
+ ## Branches and merge rules
41
+
42
+ ```
43
+ feature → develop → staging → main
44
+ ```
45
+
46
+ - Feature into the integration branch: **squash**.
47
+ - Every promotion between long-lived branches: **merge commit, never squash.** Squashing
48
+ a promotion makes the two branches look like they never met, and every promotion
49
+ afterwards shows the whole diff with conflicts, permanently.
50
+ - Protected branches: …
51
+
52
+ ---
53
+
54
+ ## Gates the project actually has
55
+
56
+ List the commands, including any that the manifest does **not** expose as a script —
57
+ those are the ones a reviewer misses.
58
+
59
+ ```bash
60
+ npm run lint
61
+ npm run typecheck
62
+ npm test
63
+ npm run build
64
+ ```
65
+
66
+ ⚠️ Note any script whose body differs from what CI runs. A checker whose npm script
67
+ omits its `--check` flag exits 0 always, and reports as a passing gate.
68
+
69
+ ---
70
+
71
+ ## Traps specific to this codebase
72
+
73
+ The highest-value section, and the only one nothing else can supply. Each entry is a
74
+ thing that has been **wrong while looking right** here.
75
+
76
+ Write each one as: *the symptom* → *the actual cause* → *what to check instead*.
77
+
78
+ - …
79
+
80
+ ---
81
+
82
+ ## What must not be published
83
+
84
+ Anything the gate should never write into a public artefact — open security items,
85
+ internal hostnames, customer identifiers.
86
+
87
+ - …