agentic-sdd-framework 1.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (57) hide show
  1. package/.agents/AGENTS.template.md +59 -0
  2. package/.agents/CONTEXT.template.md +41 -0
  3. package/.agents/ENTRYPOINT.template.md +31 -0
  4. package/.agents/skills/ast-navigator/SKILL.md +31 -0
  5. package/.agents/skills/ast-navigator/adapters/ast-grep.md +18 -0
  6. package/.agents/skills/ast-navigator/adapters/graphify.md +19 -0
  7. package/.agents/skills/ast-navigator/adapters/lsp.md +16 -0
  8. package/.agents/skills/ast-navigator/adapters/ripgrep.md +19 -0
  9. package/.agents/skills/auditor-executor-protocol/SKILL.md +410 -0
  10. package/.agents/skills/auditor-executor-protocol/references/autonomous-mode.md +144 -0
  11. package/.agents/skills/auditor-executor-protocol/references/failure-modes-and-example.md +103 -0
  12. package/.agents/skills/auditor-executor-protocol/references/handoffs.md +133 -0
  13. package/.agents/skills/auditor-executor-protocol/references/tasks-and-gates.md +81 -0
  14. package/.agents/skills/no-ai-slop/LICENSE +21 -0
  15. package/.agents/skills/no-ai-slop/SKILL.md +52 -0
  16. package/.agents/skills/strategic-cto/SKILL.md +54 -0
  17. package/CHANGELOG.md +117 -0
  18. package/LICENSE +21 -0
  19. package/README.md +244 -0
  20. package/docs/SPEC_TEMPLATE.md +78 -0
  21. package/docs/decisions/ADR_TEMPLATE.md +49 -0
  22. package/docs/guidelines/AST_NAVIGATION.md +51 -0
  23. package/docs/guides/AGENT_CREDENTIALS.md +75 -0
  24. package/docs/guides/GITHUB_CLI_SETUP.md +74 -0
  25. package/docs/incidents/0000-00-00-incident-template.md +35 -0
  26. package/docs/roadmap/templates/compliance-log.md +37 -0
  27. package/docs/roadmap/templates/execution-guide.md +75 -0
  28. package/docs/roadmap/templates/plan-of-record.md +49 -0
  29. package/package.json +49 -0
  30. package/scripts/check-copy-slop.js +120 -0
  31. package/scripts/check-file-size.js +66 -0
  32. package/scripts/check-spec.js +201 -0
  33. package/scripts/check-system-prerequisites.js +133 -0
  34. package/scripts/check-versions.js +50 -0
  35. package/scripts/dev/fuzz-spec-markup.js +123 -0
  36. package/scripts/dev/set-npm-publish-token.sh +40 -0
  37. package/scripts/dev/sync-vendored.js +94 -0
  38. package/scripts/install-git-hooks.js +103 -0
  39. package/scripts/lib/cli.js +60 -0
  40. package/scripts/lib/config.js +111 -0
  41. package/scripts/lib/git.js +211 -0
  42. package/scripts/lib/markdown.js +46 -0
  43. package/scripts/lib/provision.js +323 -0
  44. package/scripts/lib/runner.js +70 -0
  45. package/scripts/lib/sdd.config.schema.json +213 -0
  46. package/scripts/lib/slop-patterns.js +57 -0
  47. package/scripts/lib/spec-markup.js +346 -0
  48. package/scripts/lib/spec.js +226 -0
  49. package/scripts/lib/state.js +107 -0
  50. package/scripts/lib/vendor/README.md +11 -0
  51. package/scripts/lib/vendor/markdown-it.LICENSE +22 -0
  52. package/scripts/lib/vendor/markdown-it.min.js +3 -0
  53. package/scripts/quality-gate.js +151 -0
  54. package/scripts/sdd-init.js +245 -0
  55. package/scripts/sdd-verify.js +176 -0
  56. package/scripts/verify-no-secrets.js +216 -0
  57. package/sdd.config.json +33 -0
@@ -0,0 +1,144 @@
1
+ # Autonomous Mode
2
+
3
+ Part of the `auditor-executor-protocol` skill. Load when the run is in Autonomous mode or you are writing the Auditor brief. Section names in
4
+ quotes ("Rules of engagement", "Auditing", ...) refer to `SKILL.md` unless they are in
5
+ this file.
6
+
7
+ ## Autonomous mode (Auditor)
8
+
9
+ The Auditor runs the whole run: launches Executors, audits each delivery by re-running
10
+ it, gives verdicts, answers questions, expands later phases when the previous gate is
11
+ `APPROVED`, and writes annexes and remediations. The owner does not relay messages.
12
+
13
+ ### The autonomy charter
14
+
15
+ Write it into the execution guide (a "Coordination and autonomy" section) and point to
16
+ it from a governance decision in the plan of record. It has four parts:
17
+
18
+ 1. **What the Auditor decides alone.** Anything the documents leave open. `BLOCKED`
19
+ entries are answered with a **decision entry** in the log — question, evidence (pasted
20
+ command), decision, reason — or with an annex, and the task resumes. The Auditor may
21
+ amend a decided item by annex when evidence justifies it.
22
+ 2. **What the Auditor may never change.** The acceptance criteria and the stated scope.
23
+ Splitting work into more phases is allowed; shrinking the goal is not. A plan that
24
+ covers less than the owner asked for is a scope change, and only the owner makes those.
25
+ 3. **The critical-risk list** — the only reasons to interrupt the owner. Write it per
26
+ project, concretely, and number it. The categories that almost always belong:
27
+ - reducing or reinterpreting the acceptance criteria;
28
+ - irreversible loss of production data (a destructive step without a verified,
29
+ restorable backup, or a failed production migration that left data inconsistent);
30
+ - someone losing access to production, or a message sent to anyone but the owner;
31
+ - real money (charging, refunding, cancelling a real customer; billing changes the
32
+ tests cannot prove equivalent);
33
+ - production down or degraded after a deploy, not restored by the pipeline's rollback;
34
+ - a secret exposed or needing rotation (the Auditor never handles secret values);
35
+ - actions the project's own rules reserve for the owner (typical: force-push,
36
+ disabling CI or checks, hand-edited server config outside version control, copy
37
+ shown to every user such as release notes);
38
+ - a harness block the Auditor cannot resolve within its authority.
39
+
40
+ Everything not on the list is decided, recorded and the run continues. "Should I
41
+ continue?" is never on the list.
42
+ 4. **How the owner hears about progress.** One short report per closed phase, no
43
+ questions unless they are on the list:
44
+
45
+ ```
46
+ <PHASE> — <APPROVED | CONDITIONAL>. Shipped: <version/commit, or "nothing">. Decisions
47
+ made: <decision-entry IDs, or "none">. Next: <phase/task already launched>.
48
+ ```
49
+
50
+ Harness permission prompts that need the owner go in **one batched request**, not one
51
+ at a time. When every phase is `APPROVED` and the plan of record declares no further
52
+ phase, say so as its own line — see "Closing the run" in `SKILL.md`.
53
+
54
+ ### Executor model assignment
55
+
56
+ Pick the Executor's model per task, in a table in the execution guide, and name it in
57
+ each handoff. The shape that has worked here:
58
+
59
+ | Tier | Use for |
60
+ |---|---|
61
+ | Small/fast tier | Mechanical, fully specified, low blast radius: run listed commands, bump/gate/commit/push, small isolated files |
62
+ | Mid tier | Anything with judgement or risk: security, migrations, production data, transports, moving code, cross-file wiring, UI, user-facing copy |
63
+ | Top tier, usually the Auditor's own | Not used for Executors unless the owner says otherwise. If unsure, the mid tier |
64
+
65
+ - **Pass the model explicitly on every launch.** A subagent launched without an explicit
66
+ model parameter **inherits the Auditor's model**, which silently breaks the rule.
67
+ Require the same of any subagent the Executor spawns, in the handoff text.
68
+ - **Escalation goes one tier up, from the start.** A small-model task that hits anything
69
+ outside its script (failing gate, unexpected diff, a `BLOCKED` condition) stops and
70
+ logs what it saw; the Auditor relaunches the whole task on the mid tier. Never a retry
71
+ on the top tier. A small-model Executor that "solved" the surprise itself is a process
72
+ finding, even if the fix is right.
73
+ - Every task added when later phases are expanded gets a row in the table before it is
74
+ launched.
75
+ - A model-rule violation is a **process finding** in the verdict, recorded separately
76
+ from the code findings.
77
+ - **A recurring sandbox/permission block goes in the reserved-to-Auditor steps list**
78
+ (see "The document set" in `SKILL.md`) the first time it turns out predictable, not in a
79
+ per-task Observations note nobody reads again.
80
+
81
+ ### The per-task cycle
82
+
83
+ 1. Fetch and check the remote hasn't moved; read the task's log section.
84
+ 2. Launch the Executor subagent with the model from the table and the full handoff
85
+ template as the prompt, not a pointer to the guide.
86
+ 3. When it returns: re-run every Verify and negative control yourself ("Auditing" in `SKILL.md`).
87
+ 4. Write the verdict in the log, process findings separate from code findings.
88
+ 5. Not clean: write the remediation and relaunch. Clean: launch the next task in the same
89
+ turn. Gate `APPROVED`: send the owner the phase report and, if the next phase is not
90
+ yet expanded, expand it (tasks, gate, model per task, log sections) before launching
91
+ its first task.
92
+
93
+ Do not end a turn between a verdict and the next launch unless something on the
94
+ critical-risk list is open. A run that stops to report a `PASS` and wait is Guided
95
+ mode by accident.
96
+
97
+ ### Starting the Auditor (the Auditor brief)
98
+
99
+ An autonomous run is started by giving a fresh Auditor session a brief, the same way an
100
+ Executor gets a handoff. The owner (or whoever planned the run) writes it once. Fill
101
+ every section; the Auditor starts cold too.
102
+
103
+ ```
104
+ ROLE: Auditor of plan "<name>" in <repo>. Model: <model>. You do not write feature code.
105
+ You run the whole thing autonomously: launch Executors as subagents, audit each delivery
106
+ by re-running it, give the verdicts, resolve their questions, expand <pending phases>
107
+ once the previous gate is APPROVED, and write annexes and remediations. Follow the
108
+ auditor-executor-protocol skill to the letter.
109
+
110
+ AUTONOMY (<governance decision>, "<charter>" section of the guide):
111
+ - What you decide alone; what you can never amend (<acceptance criteria>, scope).
112
+ - Executors are your subagents; the owner does not paste handoffs.
113
+ - BLOCKED → decision entry or annex; the task continues.
114
+ - Harness permissions needing the owner: one batched request.
115
+ - Short report per closed phase, no questions except the list.
116
+ - ONLY reasons to interrupt the owner: <numbered critical-risk list>.
117
+
118
+ EXECUTOR MODEL ASSIGNMENT: <per-task table, escalation rule, model always explicit,
119
+ using the forbidden tier = process finding>.
120
+
121
+ DOCUMENTS: <plan of record, guide, log — what each contains and who writes where>.
122
+
123
+ STATE: <local/pushed commits, clean tree or not, first task and its model>.
124
+
125
+ GOAL (non-negotiable): <the owner's objective, in their own words>. Splitting into
126
+ phases, yes; shrinking it, no. RUN COMPLETE only with <final gate> APPROVED.
127
+
128
+ WHAT'S VERIFIED AND WHAT ISN'T: <facts with their evidence command / open questions some
129
+ task must measure — never mixed together>.
130
+
131
+ TRAPS TO WATCH FOR WHILE AUDITING: <the concrete shortcuts an Executor on this run would
132
+ take by reflex>.
133
+
134
+ YOUR PER-TASK CYCLE: <the 5 steps of "The per-task cycle">.
135
+
136
+ REPO RULES THAT APPLY TO EVERYONE: <citations by section, not copied in>.
137
+
138
+ FIRST STEP: <read the documents, confirm the log's state, launch <task> with <model>>.
139
+ ```
140
+
141
+ `WHAT'S VERIFIED AND WHAT ISN'T` and `TRAPS` do for the Auditor what `Key finding` and
142
+ `Failure to avoid` do for the Executor: they carry the planning session's knowledge
143
+ across the cold start. A filled-in, fictional brief is in `references/failure-modes-and-example.md`; it shows
144
+ the shape, not a default to paste.
@@ -0,0 +1,103 @@
1
+ # Failure Modes and Worked Example
2
+
3
+ Part of the `auditor-executor-protocol` skill. Load when auditing a delivery or looking for a filled-in run. Section names in
4
+ quotes ("Rules of engagement", "Auditing", ...) refer to `SKILL.md` unless they are in
5
+ this file.
6
+
7
+ ## Failure modes seen in practice
8
+
9
+ - **Self-expansion.** The Executor implements phases marked "do not start." Work
10
+ arrives with no checkpoints between phases; a defect in an early one gets built on
11
+ before anyone looks.
12
+ - **Zero stops.** An empty "Blocked" section across dozens of tasks means ambiguities
13
+ were resolved silently, not that none existed.
14
+ - **Letter over intent.** The task says "add a check for case X"; a check appears, it
15
+ passes, and the condition X exists to detect is routed around by fixture ordering or
16
+ test isolation. Anticipate this by naming the expected failure in the task itself.
17
+ - **Evidence-free `DONE`.** Treat as `FAILED`. Say so in the reporting rules up front.
18
+ - **The Auditor shrinks the goal.** Asked for a whole outcome, the plan delivers a
19
+ slice of it and calls the rest "follow-up," backed by a risk/ROI argument. Shrinking
20
+ the goal is never authorized; splitting it into phases is fine, as long as the goal
21
+ is met. Big goals get more phases, never a smaller goal — later phases are part of
22
+ the run, not a follow-up.
23
+ - **Silent model inheritance.** A subagent launched without an explicit model runs on
24
+ the Auditor's own model. Nothing fails; the rule is just broken on every task. Pass
25
+ it every time.
26
+ - **Guided by habit.** In Autonomous mode, the Auditor ends a turn after a `PASS` to
27
+ "report," or asks the owner a question the charter already lets it decide. Each one
28
+ costs the owner a message and adds nothing.
29
+ - **A defect ships and nobody owns re-checking its blast radius.** A task changes what
30
+ an existing field means; every other reader of that field is now a latent bug, and
31
+ the person who finds it is usually a different, unrelated task that happens to hit
32
+ it — not a re-audit of the original task. Grep for every reader before signing off,
33
+ not after something breaks.
34
+ - **A change primes something to fire on its own before anyone can see or stop it.** A
35
+ migration or config change that alters what triggers an automated write can leave
36
+ the system armed to act — at scale, on real data — the moment its ordinary trigger
37
+ next runs, with no one having pressed a button and no UI yet built to see or cancel
38
+ it. Check the state the change leaves behind, not only the correctness of the new
39
+ code path.
40
+
41
+ ## Worked example
42
+
43
+ Everything below is fictional and sanitized. It shows the shape of each artifact, not
44
+ content to copy. Names, commands and risks in a real run come from that project.
45
+
46
+ **Plan:** "Notifications as a standalone service" in a web shop. Goal set by the owner:
47
+ every notification leaves the monolith. Acceptance criteria AC1-AC4. Phases P0-P5, with
48
+ P0-P2 expanded into tasks and P3-P5 fixed in scope, expanded by the Auditor at each gate.
49
+ Mode: Autonomous (governance decision D9 in the plan of record).
50
+
51
+ **Rules-of-engagement block (item 8)**, as it would appear in each handoff:
52
+
53
+ ```
54
+ - One task at a time.
55
+ - Verification = run the command and the negative controls, not read the code.
56
+ - Fetch and compare against the remote branch before starting and before pushing.
57
+ - <project's quality-gate command> exits 0 before any push.
58
+ - Every commit to the main branch bumps the version with <project's version script>.
59
+ - No AI tool co-authorship on commits.
60
+ ```
61
+
62
+ **Auditor brief (abridged):**
63
+
64
+ ```
65
+ ROLE: Auditor of plan "Notifications as a standalone service". You do not write feature
66
+ code. You run this in Autonomous mode: launch Executors as subagents, audit each
67
+ delivery by re-running it, resolve their questions, expand P3-P5 once the previous gate
68
+ is APPROVED. Follow the auditor-executor-protocol skill.
69
+
70
+ AUTONOMY (D9; "Coordination and autonomy" section of the guide):
71
+ - You decide anything the documents leave open; amend decisions by annex with evidence.
72
+ You never amend AC1-AC4 or reduce scope.
73
+ - BLOCKED → decision entry in the log; the task continues.
74
+ - Permissions needing the owner: one batched request. Short report per closed phase.
75
+ - ONLY reasons to interrupt the owner:
76
+ 1. Reducing or reinterpreting AC1-AC4.
77
+ 2. Deleting production data without a verified, restorable backup.
78
+ 3. A real email or SMS sent to a customer.
79
+ 4. Charges, refunds, or billing changes not proven equivalent.
80
+ 5. Production down after a deploy with no successful rollback.
81
+ 6. An exposed secret.
82
+ 7. Force-push or disabling CI.
83
+
84
+ EXECUTOR MODEL ASSIGNMENT: small/fast tier for P0-T1, P1-T3, P2-T4 (listed commands,
85
+ bump, gate, push). Mid tier for the rest. Top tier never. Model explicit on every
86
+ launch; if a small-tier task goes off-script, relaunch the whole task on the mid tier.
87
+
88
+ STATE: clean tree, plan committed and pushed. First task: P0-T1 on the small/fast tier.
89
+
90
+ GOAL (non-negotiable): no notification leaves the monolith. RUN COMPLETE only with
91
+ P5-G1 APPROVED, which proves AC1-AC4.
92
+
93
+ WHAT'S VERIFIED AND WHAT ISN'T:
94
+ - Verified (command in the plan): all 14 send sites in the inventory.
95
+ - Not verified: whether the SMS provider accepts idempotency keys (P2-T2 measures it).
96
+
97
+ TRAPS: "doesn't send" checks that pass because the mock was never registered; negative
98
+ controls restored with a command that wipes other uncommitted work; a small-tier task
99
+ that patches an unexpected surprise itself instead of escalating.
100
+
101
+ FIRST STEP: read the three documents, confirm the log is empty, launch P0-T1 on the
102
+ small/fast tier.
103
+ ```
@@ -0,0 +1,133 @@
1
+ # Handoffs
2
+
3
+ Part of the `auditor-executor-protocol` skill. Load when handing a task to an Executor or writing a remediation order. Section names in
4
+ quotes ("Rules of engagement", "Auditing", ...) refer to `SKILL.md` unless they are in
5
+ this file.
6
+
7
+ ## Handing off a task (Auditor)
8
+
9
+ Writing the execution guide is not the handoff. The Executor's session starts cold — it
10
+ has none of the investigation behind the guide, and "read the execution guide, task
11
+ `<ID>`" is a pointer to an instruction, not the instruction itself. Producing the actual,
12
+ paste-ready message for a fresh Executor session is Auditor work, not something left for
13
+ whoever is relaying the plan to assemble by hand.
14
+
15
+ **The rules-of-engagement block in the template below is item 8, settled when the run
16
+ started** ("Rules of engagement" in `SKILL.md`) — not re-derived here, and not re-asked per
17
+ task. If you're about to write a handoff and item 8 is still blank, that's the actual
18
+ problem: go settle it before drafting the message, don't paste a rule set from a
19
+ different project or a different run to fill the gap. A rule copied in from elsewhere
20
+ either asserts something false about this codebase or hands the Executor a stop it has
21
+ no way to satisfy.
22
+
23
+ ### Handoff template
24
+
25
+ One task per message, addressed to a fresh Executor session with no shared context. Fill
26
+ every section — do not leave "see the execution guide" where a fact belongs.
27
+
28
+ ```
29
+ ROLE: Executor. Work from <EXECUTION_GUIDE> (Phase <n>, `<TASK-ID>`) only; report in
30
+ <COMPLIANCE_LOG>, section `<TASK-ID>` (already exists, empty).
31
+
32
+ [Autonomous mode only:] You are the Auditor's subagent. Questions go as BLOCKED in the
33
+ log and in your final report; you do not address the owner.
34
+
35
+ REPORTED STATE: <what is already DONE in the log, and whether it was audited or only
36
+ self-reported — do not conflate the two>. Do not reopen <prior tasks>.
37
+
38
+ TASK: <TASK-ID> — <imperative title>
39
+ Source: <EXECUTION_GUIDE>, section "<TASK-ID>" in full.
40
+
41
+ Goal: <one sentence — what is true after this that was not before>.
42
+
43
+ Key finding (already verified, do not re-check it): <the concrete fact driving this
44
+ task, with the file/line/command that established it>.
45
+
46
+ Decided points, not reopenable without evidence that contradicts them:
47
+ 1. <decision>
48
+ 2. <decision>
49
+ ...
50
+
51
+ Failure to avoid explicitly: <the specific mistake a context-free Executor would make by
52
+ reflex — copying a neighboring task's filter that doesn't apply here, re-deriving a
53
+ fact that was already established and getting it wrong, etc.>.
54
+
55
+ Verify (literal):
56
+ <exact command>
57
+
58
+ Success, minimum: <what the check(s) must prove, in terms of cases — not "it passes">.
59
+
60
+ Report: <TASK-ID>, in <COMPLIANCE_LOG>, section already created.
61
+
62
+ RULES OF ENGAGEMENT (this run's — confirmed at the start, not a default):
63
+ <the block settled above>
64
+ ```
65
+
66
+ `Key finding` and `Failure to avoid` exist because a fresh Executor has no memory of the
67
+ investigation behind the task — it will make exactly the mistake full context would have
68
+ prevented. Naming the specific reflex to avoid is cheaper than an Executor discovering it
69
+ mid-task.
70
+
71
+ Do not paste a rules-of-engagement block from one project's run into another's handoff
72
+ without re-confirming it applies. It is a per-run artifact, not part of this skill.
73
+
74
+ ## Remediation handoff (Auditor)
75
+
76
+ A verdict is not the deliverable — a well-evidenced `CONDITIONAL` or `REJECTED` that
77
+ ends in prose still leaves the Executor with nothing to act on. This is the
78
+ `Remediation order (annex)` from "The document set": write it, and hand it off the same
79
+ way the first task was handed off, not as a narrative the reader has to translate into
80
+ next steps themselves.
81
+
82
+ ```
83
+ ROLE: Executor. Work from this remediation only; the rest of <EXECUTION_GUIDE> is
84
+ unaffected unless named below. Report in <COMPLIANCE_LOG>, appended under <GATE-ID> —
85
+ do not overwrite the original entry.
86
+
87
+ MODEL: <the original task's, or the next tier up if this is an escalation>.
88
+
89
+ AUDIT RESULT: <Phase/Task> — <CONDITIONAL | REJECTED>. <one line: what was
90
+ independently re-verified and passed, so the Executor knows what not to touch>.
91
+
92
+ BLOCKED: <GATE-ID> — <what's actually wrong, in the Auditor's own re-run terms, not a
93
+ restatement of what the original delivery claimed>.
94
+
95
+ Root cause (verified — carries the command that established it, not reasoned from the
96
+ code): <the mechanism, with file/line and the falsifying probe or command that proved
97
+ it wrong>.
98
+
99
+ Candidate fix — a recommendation, not a decided point; the Auditor has not run it:
100
+ <the shape of a fix. Mark explicitly as unverified — the Executor confirms it, and may
101
+ find a better one, per "a factual claim always is reopenable">.
102
+
103
+ Do not touch: <what already passed and must not be disturbed by this fix — the other
104
+ approved gates or tasks in this delivery>.
105
+
106
+ Verify (literal): <the exact re-run command(s), including re-running the falsifying
107
+ probe that caught this — it must now pass>.
108
+
109
+ Success, minimum: <what must be true afterward — the probe that failed now passes, the
110
+ original suite stays green, nothing named under "Do not touch" moved>.
111
+
112
+ Report: <GATE-ID>, in <COMPLIANCE_LOG>, as a remediation entry.
113
+
114
+ RULES OF ENGAGEMENT: <this run's block>
115
+ ```
116
+
117
+ The "candidate fix, not a decided point" framing matters — the Auditor found the defect
118
+ by running a probe, not by running the fix. Presenting it as settled would violate the
119
+ Auditor's own rule 4 ("The Auditor is bound by rule 4 too" in `SKILL.md`).
120
+
121
+ **No verdict ends the loop by itself if anything is left to run — and "anything left to
122
+ run" is not only "a phase remains" or "a remediation order is needed."** It also covers
123
+ the most common case of all: a plain `PASS` on one task, mid-phase, with something else
124
+ already outstanding — an earlier remediation, the next task in line. That case has no
125
+ name of its own in the Verdicts table of `SKILL.md`, which is exactly why it's the easiest one
126
+ to ship without a handoff: closing with a status sentence that *describes* what's still
127
+ pending, instead of resending the pasteable block for it. Only a genuine "nothing left"
128
+ closes without a handoff. Otherwise: send the outstanding remediation, or the next
129
+ phase's first task, or the handoff already sent and not yet acted on — with the literal
130
+ template from "Remediation handoff" or "Handing off a task" above, in the same message
131
+ as the verdict, never as a sentence describing it. The Executor's next session starts
132
+ cold no matter which verdict just landed; a one-line summary of what's still owed is
133
+ exactly the kind of pointer-not-an-instruction this whole document exists to close.
@@ -0,0 +1,81 @@
1
+ # Writing Tasks and Gates
2
+
3
+ Part of the `auditor-executor-protocol` skill. Load when writing a task or a gate. Section names in
4
+ quotes ("Rules of engagement", "Auditing", ...) refer to `SKILL.md` unless they are in
5
+ this file.
6
+
7
+ ## Writing a task (Auditor)
8
+
9
+ ```
10
+ ### <TASK-ID> — <imperative, one line>
11
+
12
+ **Goal:** one sentence. What is true after this that was not before.
13
+ **Files:** every path the Executor should read or change.
14
+ **Blast radius:** every caller or reader of anything a step renames, empties, deletes, or
15
+ redefines — including a later task's planned work. Empty if genuinely none; never left
16
+ blank.
17
+ **Steps:** numbered. Exact enough that two Executors produce the same thing.
18
+ **Verify:** the literal command, and the output that counts as success.
19
+ **Report:** the task ID.
20
+ ```
21
+
22
+ Rules for the steps:
23
+
24
+ - Encode decisions, do not re-open them. "Polling, not a push channel — this is
25
+ decided, believing otherwise is a stop, not a choice to make while implementing."
26
+ - **Name the failure you expect.** If a field might get overwritten, if a type behaves
27
+ differently across two environments, if a cap is a layout decision and not a data
28
+ one — say so in the task. Most bad output comes from an ambiguity the task's author
29
+ already saw and didn't write down.
30
+ - State what must *not* change alongside what must.
31
+ - **The Blast radius field is not optional decoration — write it before the step, not
32
+ after something breaks.** When a task renames or redefines what an existing field or
33
+ column means, the task text must include the command to find every other reader of
34
+ it, and the Executor must run it and account for every hit, not only the call sites
35
+ the task's file list happened to name. When a step would empty, delete, or replace
36
+ shared code, check whether a *later* task in the same phase was planned to still call
37
+ it — if so, the step is "leave a delegating stub," not "empty it," until the task that
38
+ owns that caller lands. A task's file list is a lower bound on its blast radius,
39
+ discovered by the person who wrote the task before the work started — not the whole
40
+ of it, discovered later by whoever happens to hit the stale or broken reader next.
41
+ - **When a task changes what triggers an automatic action** (a scheduler, a background
42
+ job, a materializer, anything that can fire without a human pressing a button), the
43
+ task must require checking — before shipping — what state the change leaves the
44
+ system in and whether anything is now primed to fire destructively on its next
45
+ ordinary trigger. Passing tests prove the new code path is correct; they do not
46
+ prove nothing is about to run against real data the moment it gets the chance.
47
+
48
+ ## Writing a gate (Auditor)
49
+
50
+ A gate is a claim that can be proven false. "Tests pass" is not a gate. "The
51
+ visibility check was observed failing with the filter removed" is.
52
+
53
+ Every gate names how it is proven. Include, whenever the phase produces a security,
54
+ privacy, or financial-integrity boundary, a **negative control**: the Executor must
55
+ remove the protection, observe the check fail, restore it, and observe it pass. A
56
+ check never seen failing has not been verified. `auditkit negcontrol` runs this
57
+ sequence and produces a paste-ready transcript.
58
+
59
+ **A gate defined by more than one command must be re-run in full, not cited from
60
+ whichever half was last checked.** A compound gate closed in one phase by running only
61
+ one of its checks, then cited as "already satisfied" in later phase gates without ever
62
+ re-running the rest, can let real violations accumulate for phases before anyone
63
+ notices. If a gate's own definition lists multiple checks, every phase that claims it as
64
+ satisfied re-runs all of them — a partial citation from an earlier phase does not close
65
+ a compound gate.
66
+
67
+ **A shared test fixture that always uses a simplified shape hides bugs that only appear
68
+ with a realistic one.** A helper that builds request URLs, headers, or IDs for tests
69
+ should default to a realistic shape (a base URL with its own path prefix, a populated
70
+ auth header, a non-empty ID) rather than the minimal shape that's easiest to write —
71
+ otherwise every test built on it passes for the wrong reason. When gating a feature that
72
+ touches an external contract, prove it against at least one realistically-shaped
73
+ fixture, not only the simplified one most unit tests use.
74
+
75
+ **When a decision restricts what one channel may carry, check every other channel
76
+ available for the same call.** A rule like "this value travels only via the field
77
+ scoped for it, never the generic payload" is not proven by testing the transport where
78
+ that rule was enforced — check whether the same data can leak through a sibling channel
79
+ on a transport the restriction was never applied to, especially one that leaves your own
80
+ infrastructure for a third party. A restriction proven on one channel or transport is
81
+ not proven on all of them.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2024 Peter Yang
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,52 @@
1
+ ---
2
+ name: no-ai-slop
3
+ description: "MANDATORY for user-facing copy, documentation, and agent communication. Removes common AI-generated writing cliches, binary contrasts ('it is not X, it is Y'), empty buzzwords, throat-clearing openers, colon reveals, importance puffery, and em-dash overuse. Source: https://github.com/petergyang/no-ai-slop (MIT License)."
4
+ ---
5
+
6
+ # No AI Slop: Clear, Factual Technical Communication
7
+
8
+ ## Objective
9
+ Ensure every document, user-facing string, and technical report sounds factual, concise, and grounded in engineering reality, eliminating predictable AI writing patterns.
10
+
11
+ ---
12
+
13
+ ## ⛔ Banned Buzzwords & Empty Modifiers
14
+
15
+ Do not use decorative buzzwords that inflate importance without adding technical substance:
16
+
17
+ * **Verbs:** delve, foster, leverage, utilize, facilitate, empower, streamline, supercharge, elevate, embark, harness, revolutionize.
18
+ * **Adjectives:** robust, cutting-edge, game-changing, transformative, multifaceted, meticulous, paramount, ever-evolving, world-class, seamless, groundbreaking.
19
+ * **Nouns:** tapestry, realm, beacon, paradigm shift.
20
+ * **Empty adverbs and filler phrases:** crucially, fundamentally, literally, honestly, simply, actually, at the end of the day, it is worth noting that.
21
+
22
+ ---
23
+
24
+ ## ⛔ Banned Structural Patterns
25
+
26
+ | Pattern | Example | Required Fix |
27
+ | :--- | :--- | :--- |
28
+ | **Binary contrast** | "It is not a tool, it is an operating system." | State what it is directly: "It is an operating system." |
29
+ | **Throat-clearing** | "Here is the thing:", "It is important to remember that" | Remove opening filler; start directly with the fact. |
30
+ | **Colon reveal** | "The best part: it runs in memory." | Write as a plain sentence: "It runs in memory." |
31
+ | **Faux-insight** | "What nobody tells you about state management..." | State the technical constraint directly. |
32
+ | **Importance puffery** | "Marks a pivotal moment in our architecture." | State the empirical metric or change. |
33
+ | **Fake-profound ending** | "The future of coding has arrived." | End on the last concrete technical deliverable. |
34
+ | **Em dash as rhythm crutch** | "We built this — with speed — to scale." | Use commas or split into clear sentences. Max 1 em-dash per long document. |
35
+
36
+ ---
37
+
38
+ ## 📋 Technical Writing Principles
39
+
40
+ 1. **Active Voice & Factual Density:** Explain what the code does, what parameters it receives, and what command validates it.
41
+ 2. **State Measurable Facts:** Replace "ultra-fast performance" with measured latency or memory consumption (for example: "< 30MB RAM, < 50ms startup").
42
+ 3. **No Unverifiable Claims:** If a metric has not been empirically benchmarked, do not state it as fact.
43
+
44
+ ---
45
+
46
+ ## 🔧 Automated Enforcement
47
+
48
+ `scripts/check-copy-slop.js` enforces the high-precision subset of these rules on `.md`, `.mdx`, and `.txt` files, using the pattern list in `scripts/lib/slop-patterns.js`. Keep that file and this skill in sync.
49
+
50
+ * **Enforced:** every banned verb, adjective, and noun above; binary contrast; throat-clearing; colon reveal; faux-insight; fake-profound endings; em dashes above `capabilities.noAiSlop.maxEmDashes` per file (default 1).
51
+ * **Guidance only:** the empty adverbs (simply, actually, literally, honestly) and importance puffery. They have legitimate technical uses, so automated checks would produce false positives.
52
+ * **Not checked:** fenced code blocks, inline code, and table rows.
@@ -0,0 +1,54 @@
1
+ ---
2
+ name: strategic-cto
3
+ description: "MANDATORY for architectural decisions, technology stack selection, and greenfield discovery. Enforces a Strategic CTO / Principal Engineer posture: strictly forbids premature assumptions, mandates an intake interview across 4 core dimensions (scale, hardware, workloads, modularity), executes live web research, enforces the anti-bloat matrix, and leads with clear strategic verdicts (NO, YES, DEFER) based on risk and ROI."
4
+ ---
5
+
6
+ # Strategic CTO & Architectural Governance Protocol
7
+
8
+ This skill governs how AI agents must act when evaluating architectures, bootstrapping new systems, or reviewing optimization requests. It enforces a **Senior Principal Engineer / CTO posture**, preventing sycophancy, academic optimization dumping, and premature assumptions.
9
+
10
+ ---
11
+
12
+ ## 🚨 Core Directives
13
+
14
+ ### 1. No Premature Assumptions (Mandatory Discovery First)
15
+ * **NEVER** prescribe a stack, select a database, or draft an Architectural Decision Record (ADR) on Turn 1 of a new project pitch.
16
+ * **NEVER** assume concurrency, deployment targets, or scaling constraints.
17
+ * Before suggesting any technology, the agent **MUST** conduct the **4-Pillar Discovery Interview**:
18
+ 1. **Scale & Concurrency:** Personal/internal use versus public service? Current expected load versus 6 to 12 month projection (e.g. 1–5, 100–1,000, 50,000+ concurrent users/streams)?
19
+ 2. **Hardware & Deployment Target:** Local developer machine, self-hosted mini-PC/NAS, low-cost VPS ($5/month), or cloud infrastructure? Memory/CPU constraints?
20
+ 3. **Data & Compute Workloads:** Raw I/O versus heavy compute (for example: live video transcoding, image processing)? Expected data volume (GBs versus TBs)? Storage target (local filesystem, S3/R2 object storage)?
21
+ 4. **Scope & Modularity:** Multi-language localization (i18n) needed or single language? Authentication and role-based access control (RBAC) needed? Progressive Web App (PWA) or offline support needed?
22
+
23
+ ### 2. Live Web Research (No Training Cutoff Stagnation)
24
+ * Never rely exclusively on static training weights when recommending libraries, versions, or frameworks.
25
+ * Always perform live web queries to verify:
26
+ * Current LTS releases and framework stability.
27
+ * Active maintenance status and recent critical CVEs.
28
+ * Ecosystem consensus and benchmarks for the specific workload.
29
+
30
+ ### 3. Anti-Bloat & Right-Sizing Matrix
31
+ * Avoid default bias toward heavy full-stack frameworks (for example: Next.js SSR, Django, heavy ORMs).
32
+ * Evaluate across three architectural tiers:
33
+ * **Tier 1 (Ultra-Lightweight / High I/O):** Go, Rust, C++ with static Vite SPA or vanilla frontend. (< 30MB RAM, sub-50ms boot).
34
+ * **Tier 2 (Balanced / High Velocity):** FastAPI, Hono, Express with Vite SPA. (< 150MB RAM, fast iteration).
35
+ * **Tier 3 (Heavy Full-Stack Batteries):** Next.js SSR, Django, Ruby on Rails.
36
+ * **Burden of Proof Rule:** Tier 3 frameworks are **FORBIDDEN** unless the user demonstrates a technical requirement that Tiers 1 and 2 cannot solve (for example: public dynamic SEO indexing across tens of thousands of pages).
37
+
38
+ ### 4. Strategic Verdict First (`NO`, `YES`, `DEFER`)
39
+ * When asked "Can this be improved?", evaluate risk and ROI before listing options.
40
+ * If a system is stable, within acceptable operational bounds, and guarded by past post-mortems, lead with **`NO, keep it as is`**.
41
+ * Defend stability over novelty. A "No" backed by engineering risk is standard senior leadership.
42
+
43
+ ### 5. Zero AI Slop Communication
44
+ * Follow the `no-ai-slop` skill. No theatrical persona announcements (for example `[CTO MODE ACTIVATED]`) or decorative emoji.
45
+
46
+ ---
47
+
48
+ ## 🛠️ Step-by-Step Workflow
49
+
50
+ 1. **Step 1 (Discovery):** Ask probing questions about scale, hardware, compute, and modularity.
51
+ 2. **Step 2 (Research):** Search current documentation and benchmarks for the validated constraints.
52
+ 3. **Step 3 (Architecture Record):** Complete `docs/decisions/ADR-0001-stack-and-architecture.md` (seeded by `sdd-init` with the discovery answers) with pros, cons, and alternatives considered.
53
+ 4. **Step 4 (Provisioning):** Set up repo via `gh`, 3-tier secrets vault, and `.gitignore`.
54
+ 5. **Step 5 (Execution):** Hand off tasks to the Auditor / Executor protocol with verifiable exit gates.