agentic-sdd-framework 1.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/AGENTS.template.md +59 -0
- package/.agents/CONTEXT.template.md +41 -0
- package/.agents/ENTRYPOINT.template.md +31 -0
- package/.agents/skills/ast-navigator/SKILL.md +31 -0
- package/.agents/skills/ast-navigator/adapters/ast-grep.md +18 -0
- package/.agents/skills/ast-navigator/adapters/graphify.md +19 -0
- package/.agents/skills/ast-navigator/adapters/lsp.md +16 -0
- package/.agents/skills/ast-navigator/adapters/ripgrep.md +19 -0
- package/.agents/skills/auditor-executor-protocol/SKILL.md +410 -0
- package/.agents/skills/auditor-executor-protocol/references/autonomous-mode.md +144 -0
- package/.agents/skills/auditor-executor-protocol/references/failure-modes-and-example.md +103 -0
- package/.agents/skills/auditor-executor-protocol/references/handoffs.md +133 -0
- package/.agents/skills/auditor-executor-protocol/references/tasks-and-gates.md +81 -0
- package/.agents/skills/no-ai-slop/LICENSE +21 -0
- package/.agents/skills/no-ai-slop/SKILL.md +52 -0
- package/.agents/skills/strategic-cto/SKILL.md +54 -0
- package/CHANGELOG.md +117 -0
- package/LICENSE +21 -0
- package/README.md +244 -0
- package/docs/SPEC_TEMPLATE.md +78 -0
- package/docs/decisions/ADR_TEMPLATE.md +49 -0
- package/docs/guidelines/AST_NAVIGATION.md +51 -0
- package/docs/guides/AGENT_CREDENTIALS.md +75 -0
- package/docs/guides/GITHUB_CLI_SETUP.md +74 -0
- package/docs/incidents/0000-00-00-incident-template.md +35 -0
- package/docs/roadmap/templates/compliance-log.md +37 -0
- package/docs/roadmap/templates/execution-guide.md +75 -0
- package/docs/roadmap/templates/plan-of-record.md +49 -0
- package/package.json +49 -0
- package/scripts/check-copy-slop.js +120 -0
- package/scripts/check-file-size.js +66 -0
- package/scripts/check-spec.js +201 -0
- package/scripts/check-system-prerequisites.js +133 -0
- package/scripts/check-versions.js +50 -0
- package/scripts/dev/fuzz-spec-markup.js +123 -0
- package/scripts/dev/set-npm-publish-token.sh +40 -0
- package/scripts/dev/sync-vendored.js +94 -0
- package/scripts/install-git-hooks.js +103 -0
- package/scripts/lib/cli.js +60 -0
- package/scripts/lib/config.js +111 -0
- package/scripts/lib/git.js +211 -0
- package/scripts/lib/markdown.js +46 -0
- package/scripts/lib/provision.js +323 -0
- package/scripts/lib/runner.js +70 -0
- package/scripts/lib/sdd.config.schema.json +213 -0
- package/scripts/lib/slop-patterns.js +57 -0
- package/scripts/lib/spec-markup.js +346 -0
- package/scripts/lib/spec.js +226 -0
- package/scripts/lib/state.js +107 -0
- package/scripts/lib/vendor/README.md +11 -0
- package/scripts/lib/vendor/markdown-it.LICENSE +22 -0
- package/scripts/lib/vendor/markdown-it.min.js +3 -0
- package/scripts/quality-gate.js +151 -0
- package/scripts/sdd-init.js +245 -0
- package/scripts/sdd-verify.js +176 -0
- package/scripts/verify-no-secrets.js +216 -0
- package/sdd.config.json +33 -0
|
@@ -0,0 +1,144 @@
|
|
|
1
|
+
# Autonomous Mode
|
|
2
|
+
|
|
3
|
+
Part of the `auditor-executor-protocol` skill. Load when the run is in Autonomous mode or you are writing the Auditor brief. Section names in
|
|
4
|
+
quotes ("Rules of engagement", "Auditing", ...) refer to `SKILL.md` unless they are in
|
|
5
|
+
this file.
|
|
6
|
+
|
|
7
|
+
## Autonomous mode (Auditor)
|
|
8
|
+
|
|
9
|
+
The Auditor runs the whole run: launches Executors, audits each delivery by re-running
|
|
10
|
+
it, gives verdicts, answers questions, expands later phases when the previous gate is
|
|
11
|
+
`APPROVED`, and writes annexes and remediations. The owner does not relay messages.
|
|
12
|
+
|
|
13
|
+
### The autonomy charter
|
|
14
|
+
|
|
15
|
+
Write it into the execution guide (a "Coordination and autonomy" section) and point to
|
|
16
|
+
it from a governance decision in the plan of record. It has four parts:
|
|
17
|
+
|
|
18
|
+
1. **What the Auditor decides alone.** Anything the documents leave open. `BLOCKED`
|
|
19
|
+
entries are answered with a **decision entry** in the log — question, evidence (pasted
|
|
20
|
+
command), decision, reason — or with an annex, and the task resumes. The Auditor may
|
|
21
|
+
amend a decided item by annex when evidence justifies it.
|
|
22
|
+
2. **What the Auditor may never change.** The acceptance criteria and the stated scope.
|
|
23
|
+
Splitting work into more phases is allowed; shrinking the goal is not. A plan that
|
|
24
|
+
covers less than the owner asked for is a scope change, and only the owner makes those.
|
|
25
|
+
3. **The critical-risk list** — the only reasons to interrupt the owner. Write it per
|
|
26
|
+
project, concretely, and number it. The categories that almost always belong:
|
|
27
|
+
- reducing or reinterpreting the acceptance criteria;
|
|
28
|
+
- irreversible loss of production data (a destructive step without a verified,
|
|
29
|
+
restorable backup, or a failed production migration that left data inconsistent);
|
|
30
|
+
- someone losing access to production, or a message sent to anyone but the owner;
|
|
31
|
+
- real money (charging, refunding, cancelling a real customer; billing changes the
|
|
32
|
+
tests cannot prove equivalent);
|
|
33
|
+
- production down or degraded after a deploy, not restored by the pipeline's rollback;
|
|
34
|
+
- a secret exposed or needing rotation (the Auditor never handles secret values);
|
|
35
|
+
- actions the project's own rules reserve for the owner (typical: force-push,
|
|
36
|
+
disabling CI or checks, hand-edited server config outside version control, copy
|
|
37
|
+
shown to every user such as release notes);
|
|
38
|
+
- a harness block the Auditor cannot resolve within its authority.
|
|
39
|
+
|
|
40
|
+
Everything not on the list is decided, recorded and the run continues. "Should I
|
|
41
|
+
continue?" is never on the list.
|
|
42
|
+
4. **How the owner hears about progress.** One short report per closed phase, no
|
|
43
|
+
questions unless they are on the list:
|
|
44
|
+
|
|
45
|
+
```
|
|
46
|
+
<PHASE> — <APPROVED | CONDITIONAL>. Shipped: <version/commit, or "nothing">. Decisions
|
|
47
|
+
made: <decision-entry IDs, or "none">. Next: <phase/task already launched>.
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
Harness permission prompts that need the owner go in **one batched request**, not one
|
|
51
|
+
at a time. When every phase is `APPROVED` and the plan of record declares no further
|
|
52
|
+
phase, say so as its own line — see "Closing the run" in `SKILL.md`.
|
|
53
|
+
|
|
54
|
+
### Executor model assignment
|
|
55
|
+
|
|
56
|
+
Pick the Executor's model per task, in a table in the execution guide, and name it in
|
|
57
|
+
each handoff. The shape that has worked here:
|
|
58
|
+
|
|
59
|
+
| Tier | Use for |
|
|
60
|
+
|---|---|
|
|
61
|
+
| Small/fast tier | Mechanical, fully specified, low blast radius: run listed commands, bump/gate/commit/push, small isolated files |
|
|
62
|
+
| Mid tier | Anything with judgement or risk: security, migrations, production data, transports, moving code, cross-file wiring, UI, user-facing copy |
|
|
63
|
+
| Top tier, usually the Auditor's own | Not used for Executors unless the owner says otherwise. If unsure, the mid tier |
|
|
64
|
+
|
|
65
|
+
- **Pass the model explicitly on every launch.** A subagent launched without an explicit
|
|
66
|
+
model parameter **inherits the Auditor's model**, which silently breaks the rule.
|
|
67
|
+
Require the same of any subagent the Executor spawns, in the handoff text.
|
|
68
|
+
- **Escalation goes one tier up, from the start.** A small-model task that hits anything
|
|
69
|
+
outside its script (failing gate, unexpected diff, a `BLOCKED` condition) stops and
|
|
70
|
+
logs what it saw; the Auditor relaunches the whole task on the mid tier. Never a retry
|
|
71
|
+
on the top tier. A small-model Executor that "solved" the surprise itself is a process
|
|
72
|
+
finding, even if the fix is right.
|
|
73
|
+
- Every task added when later phases are expanded gets a row in the table before it is
|
|
74
|
+
launched.
|
|
75
|
+
- A model-rule violation is a **process finding** in the verdict, recorded separately
|
|
76
|
+
from the code findings.
|
|
77
|
+
- **A recurring sandbox/permission block goes in the reserved-to-Auditor steps list**
|
|
78
|
+
(see "The document set" in `SKILL.md`) the first time it turns out predictable, not in a
|
|
79
|
+
per-task Observations note nobody reads again.
|
|
80
|
+
|
|
81
|
+
### The per-task cycle
|
|
82
|
+
|
|
83
|
+
1. Fetch and check the remote hasn't moved; read the task's log section.
|
|
84
|
+
2. Launch the Executor subagent with the model from the table and the full handoff
|
|
85
|
+
template as the prompt, not a pointer to the guide.
|
|
86
|
+
3. When it returns: re-run every Verify and negative control yourself ("Auditing" in `SKILL.md`).
|
|
87
|
+
4. Write the verdict in the log, process findings separate from code findings.
|
|
88
|
+
5. Not clean: write the remediation and relaunch. Clean: launch the next task in the same
|
|
89
|
+
turn. Gate `APPROVED`: send the owner the phase report and, if the next phase is not
|
|
90
|
+
yet expanded, expand it (tasks, gate, model per task, log sections) before launching
|
|
91
|
+
its first task.
|
|
92
|
+
|
|
93
|
+
Do not end a turn between a verdict and the next launch unless something on the
|
|
94
|
+
critical-risk list is open. A run that stops to report a `PASS` and wait is Guided
|
|
95
|
+
mode by accident.
|
|
96
|
+
|
|
97
|
+
### Starting the Auditor (the Auditor brief)
|
|
98
|
+
|
|
99
|
+
An autonomous run is started by giving a fresh Auditor session a brief, the same way an
|
|
100
|
+
Executor gets a handoff. The owner (or whoever planned the run) writes it once. Fill
|
|
101
|
+
every section; the Auditor starts cold too.
|
|
102
|
+
|
|
103
|
+
```
|
|
104
|
+
ROLE: Auditor of plan "<name>" in <repo>. Model: <model>. You do not write feature code.
|
|
105
|
+
You run the whole thing autonomously: launch Executors as subagents, audit each delivery
|
|
106
|
+
by re-running it, give the verdicts, resolve their questions, expand <pending phases>
|
|
107
|
+
once the previous gate is APPROVED, and write annexes and remediations. Follow the
|
|
108
|
+
auditor-executor-protocol skill to the letter.
|
|
109
|
+
|
|
110
|
+
AUTONOMY (<governance decision>, "<charter>" section of the guide):
|
|
111
|
+
- What you decide alone; what you can never amend (<acceptance criteria>, scope).
|
|
112
|
+
- Executors are your subagents; the owner does not paste handoffs.
|
|
113
|
+
- BLOCKED → decision entry or annex; the task continues.
|
|
114
|
+
- Harness permissions needing the owner: one batched request.
|
|
115
|
+
- Short report per closed phase, no questions except the list.
|
|
116
|
+
- ONLY reasons to interrupt the owner: <numbered critical-risk list>.
|
|
117
|
+
|
|
118
|
+
EXECUTOR MODEL ASSIGNMENT: <per-task table, escalation rule, model always explicit,
|
|
119
|
+
using the forbidden tier = process finding>.
|
|
120
|
+
|
|
121
|
+
DOCUMENTS: <plan of record, guide, log — what each contains and who writes where>.
|
|
122
|
+
|
|
123
|
+
STATE: <local/pushed commits, clean tree or not, first task and its model>.
|
|
124
|
+
|
|
125
|
+
GOAL (non-negotiable): <the owner's objective, in their own words>. Splitting into
|
|
126
|
+
phases, yes; shrinking it, no. RUN COMPLETE only with <final gate> APPROVED.
|
|
127
|
+
|
|
128
|
+
WHAT'S VERIFIED AND WHAT ISN'T: <facts with their evidence command / open questions some
|
|
129
|
+
task must measure — never mixed together>.
|
|
130
|
+
|
|
131
|
+
TRAPS TO WATCH FOR WHILE AUDITING: <the concrete shortcuts an Executor on this run would
|
|
132
|
+
take by reflex>.
|
|
133
|
+
|
|
134
|
+
YOUR PER-TASK CYCLE: <the 5 steps of "The per-task cycle">.
|
|
135
|
+
|
|
136
|
+
REPO RULES THAT APPLY TO EVERYONE: <citations by section, not copied in>.
|
|
137
|
+
|
|
138
|
+
FIRST STEP: <read the documents, confirm the log's state, launch <task> with <model>>.
|
|
139
|
+
```
|
|
140
|
+
|
|
141
|
+
`WHAT'S VERIFIED AND WHAT ISN'T` and `TRAPS` do for the Auditor what `Key finding` and
|
|
142
|
+
`Failure to avoid` do for the Executor: they carry the planning session's knowledge
|
|
143
|
+
across the cold start. A filled-in, fictional brief is in `references/failure-modes-and-example.md`; it shows
|
|
144
|
+
the shape, not a default to paste.
|
|
@@ -0,0 +1,103 @@
|
|
|
1
|
+
# Failure Modes and Worked Example
|
|
2
|
+
|
|
3
|
+
Part of the `auditor-executor-protocol` skill. Load when auditing a delivery or looking for a filled-in run. Section names in
|
|
4
|
+
quotes ("Rules of engagement", "Auditing", ...) refer to `SKILL.md` unless they are in
|
|
5
|
+
this file.
|
|
6
|
+
|
|
7
|
+
## Failure modes seen in practice
|
|
8
|
+
|
|
9
|
+
- **Self-expansion.** The Executor implements phases marked "do not start." Work
|
|
10
|
+
arrives with no checkpoints between phases; a defect in an early one gets built on
|
|
11
|
+
before anyone looks.
|
|
12
|
+
- **Zero stops.** An empty "Blocked" section across dozens of tasks means ambiguities
|
|
13
|
+
were resolved silently, not that none existed.
|
|
14
|
+
- **Letter over intent.** The task says "add a check for case X"; a check appears, it
|
|
15
|
+
passes, and the condition X exists to detect is routed around by fixture ordering or
|
|
16
|
+
test isolation. Anticipate this by naming the expected failure in the task itself.
|
|
17
|
+
- **Evidence-free `DONE`.** Treat as `FAILED`. Say so in the reporting rules up front.
|
|
18
|
+
- **The Auditor shrinks the goal.** Asked for a whole outcome, the plan delivers a
|
|
19
|
+
slice of it and calls the rest "follow-up," backed by a risk/ROI argument. Shrinking
|
|
20
|
+
the goal is never authorized; splitting it into phases is fine, as long as the goal
|
|
21
|
+
is met. Big goals get more phases, never a smaller goal — later phases are part of
|
|
22
|
+
the run, not a follow-up.
|
|
23
|
+
- **Silent model inheritance.** A subagent launched without an explicit model runs on
|
|
24
|
+
the Auditor's own model. Nothing fails; the rule is just broken on every task. Pass
|
|
25
|
+
it every time.
|
|
26
|
+
- **Guided by habit.** In Autonomous mode, the Auditor ends a turn after a `PASS` to
|
|
27
|
+
"report," or asks the owner a question the charter already lets it decide. Each one
|
|
28
|
+
costs the owner a message and adds nothing.
|
|
29
|
+
- **A defect ships and nobody owns re-checking its blast radius.** A task changes what
|
|
30
|
+
an existing field means; every other reader of that field is now a latent bug, and
|
|
31
|
+
the person who finds it is usually a different, unrelated task that happens to hit
|
|
32
|
+
it — not a re-audit of the original task. Grep for every reader before signing off,
|
|
33
|
+
not after something breaks.
|
|
34
|
+
- **A change primes something to fire on its own before anyone can see or stop it.** A
|
|
35
|
+
migration or config change that alters what triggers an automated write can leave
|
|
36
|
+
the system armed to act — at scale, on real data — the moment its ordinary trigger
|
|
37
|
+
next runs, with no one having pressed a button and no UI yet built to see or cancel
|
|
38
|
+
it. Check the state the change leaves behind, not only the correctness of the new
|
|
39
|
+
code path.
|
|
40
|
+
|
|
41
|
+
## Worked example
|
|
42
|
+
|
|
43
|
+
Everything below is fictional and sanitized. It shows the shape of each artifact, not
|
|
44
|
+
content to copy. Names, commands and risks in a real run come from that project.
|
|
45
|
+
|
|
46
|
+
**Plan:** "Notifications as a standalone service" in a web shop. Goal set by the owner:
|
|
47
|
+
every notification leaves the monolith. Acceptance criteria AC1-AC4. Phases P0-P5, with
|
|
48
|
+
P0-P2 expanded into tasks and P3-P5 fixed in scope, expanded by the Auditor at each gate.
|
|
49
|
+
Mode: Autonomous (governance decision D9 in the plan of record).
|
|
50
|
+
|
|
51
|
+
**Rules-of-engagement block (item 8)**, as it would appear in each handoff:
|
|
52
|
+
|
|
53
|
+
```
|
|
54
|
+
- One task at a time.
|
|
55
|
+
- Verification = run the command and the negative controls, not read the code.
|
|
56
|
+
- Fetch and compare against the remote branch before starting and before pushing.
|
|
57
|
+
- <project's quality-gate command> exits 0 before any push.
|
|
58
|
+
- Every commit to the main branch bumps the version with <project's version script>.
|
|
59
|
+
- No AI tool co-authorship on commits.
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
**Auditor brief (abridged):**
|
|
63
|
+
|
|
64
|
+
```
|
|
65
|
+
ROLE: Auditor of plan "Notifications as a standalone service". You do not write feature
|
|
66
|
+
code. You run this in Autonomous mode: launch Executors as subagents, audit each
|
|
67
|
+
delivery by re-running it, resolve their questions, expand P3-P5 once the previous gate
|
|
68
|
+
is APPROVED. Follow the auditor-executor-protocol skill.
|
|
69
|
+
|
|
70
|
+
AUTONOMY (D9; "Coordination and autonomy" section of the guide):
|
|
71
|
+
- You decide anything the documents leave open; amend decisions by annex with evidence.
|
|
72
|
+
You never amend AC1-AC4 or reduce scope.
|
|
73
|
+
- BLOCKED → decision entry in the log; the task continues.
|
|
74
|
+
- Permissions needing the owner: one batched request. Short report per closed phase.
|
|
75
|
+
- ONLY reasons to interrupt the owner:
|
|
76
|
+
1. Reducing or reinterpreting AC1-AC4.
|
|
77
|
+
2. Deleting production data without a verified, restorable backup.
|
|
78
|
+
3. A real email or SMS sent to a customer.
|
|
79
|
+
4. Charges, refunds, or billing changes not proven equivalent.
|
|
80
|
+
5. Production down after a deploy with no successful rollback.
|
|
81
|
+
6. An exposed secret.
|
|
82
|
+
7. Force-push or disabling CI.
|
|
83
|
+
|
|
84
|
+
EXECUTOR MODEL ASSIGNMENT: small/fast tier for P0-T1, P1-T3, P2-T4 (listed commands,
|
|
85
|
+
bump, gate, push). Mid tier for the rest. Top tier never. Model explicit on every
|
|
86
|
+
launch; if a small-tier task goes off-script, relaunch the whole task on the mid tier.
|
|
87
|
+
|
|
88
|
+
STATE: clean tree, plan committed and pushed. First task: P0-T1 on the small/fast tier.
|
|
89
|
+
|
|
90
|
+
GOAL (non-negotiable): no notification leaves the monolith. RUN COMPLETE only with
|
|
91
|
+
P5-G1 APPROVED, which proves AC1-AC4.
|
|
92
|
+
|
|
93
|
+
WHAT'S VERIFIED AND WHAT ISN'T:
|
|
94
|
+
- Verified (command in the plan): all 14 send sites in the inventory.
|
|
95
|
+
- Not verified: whether the SMS provider accepts idempotency keys (P2-T2 measures it).
|
|
96
|
+
|
|
97
|
+
TRAPS: "doesn't send" checks that pass because the mock was never registered; negative
|
|
98
|
+
controls restored with a command that wipes other uncommitted work; a small-tier task
|
|
99
|
+
that patches an unexpected surprise itself instead of escalating.
|
|
100
|
+
|
|
101
|
+
FIRST STEP: read the three documents, confirm the log is empty, launch P0-T1 on the
|
|
102
|
+
small/fast tier.
|
|
103
|
+
```
|
|
@@ -0,0 +1,133 @@
|
|
|
1
|
+
# Handoffs
|
|
2
|
+
|
|
3
|
+
Part of the `auditor-executor-protocol` skill. Load when handing a task to an Executor or writing a remediation order. Section names in
|
|
4
|
+
quotes ("Rules of engagement", "Auditing", ...) refer to `SKILL.md` unless they are in
|
|
5
|
+
this file.
|
|
6
|
+
|
|
7
|
+
## Handing off a task (Auditor)
|
|
8
|
+
|
|
9
|
+
Writing the execution guide is not the handoff. The Executor's session starts cold — it
|
|
10
|
+
has none of the investigation behind the guide, and "read the execution guide, task
|
|
11
|
+
`<ID>`" is a pointer to an instruction, not the instruction itself. Producing the actual,
|
|
12
|
+
paste-ready message for a fresh Executor session is Auditor work, not something left for
|
|
13
|
+
whoever is relaying the plan to assemble by hand.
|
|
14
|
+
|
|
15
|
+
**The rules-of-engagement block in the template below is item 8, settled when the run
|
|
16
|
+
started** ("Rules of engagement" in `SKILL.md`) — not re-derived here, and not re-asked per
|
|
17
|
+
task. If you're about to write a handoff and item 8 is still blank, that's the actual
|
|
18
|
+
problem: go settle it before drafting the message, don't paste a rule set from a
|
|
19
|
+
different project or a different run to fill the gap. A rule copied in from elsewhere
|
|
20
|
+
either asserts something false about this codebase or hands the Executor a stop it has
|
|
21
|
+
no way to satisfy.
|
|
22
|
+
|
|
23
|
+
### Handoff template
|
|
24
|
+
|
|
25
|
+
One task per message, addressed to a fresh Executor session with no shared context. Fill
|
|
26
|
+
every section — do not leave "see the execution guide" where a fact belongs.
|
|
27
|
+
|
|
28
|
+
```
|
|
29
|
+
ROLE: Executor. Work from <EXECUTION_GUIDE> (Phase <n>, `<TASK-ID>`) only; report in
|
|
30
|
+
<COMPLIANCE_LOG>, section `<TASK-ID>` (already exists, empty).
|
|
31
|
+
|
|
32
|
+
[Autonomous mode only:] You are the Auditor's subagent. Questions go as BLOCKED in the
|
|
33
|
+
log and in your final report; you do not address the owner.
|
|
34
|
+
|
|
35
|
+
REPORTED STATE: <what is already DONE in the log, and whether it was audited or only
|
|
36
|
+
self-reported — do not conflate the two>. Do not reopen <prior tasks>.
|
|
37
|
+
|
|
38
|
+
TASK: <TASK-ID> — <imperative title>
|
|
39
|
+
Source: <EXECUTION_GUIDE>, section "<TASK-ID>" in full.
|
|
40
|
+
|
|
41
|
+
Goal: <one sentence — what is true after this that was not before>.
|
|
42
|
+
|
|
43
|
+
Key finding (already verified, do not re-check it): <the concrete fact driving this
|
|
44
|
+
task, with the file/line/command that established it>.
|
|
45
|
+
|
|
46
|
+
Decided points, not reopenable without evidence that contradicts them:
|
|
47
|
+
1. <decision>
|
|
48
|
+
2. <decision>
|
|
49
|
+
...
|
|
50
|
+
|
|
51
|
+
Failure to avoid explicitly: <the specific mistake a context-free Executor would make by
|
|
52
|
+
reflex — copying a neighboring task's filter that doesn't apply here, re-deriving a
|
|
53
|
+
fact that was already established and getting it wrong, etc.>.
|
|
54
|
+
|
|
55
|
+
Verify (literal):
|
|
56
|
+
<exact command>
|
|
57
|
+
|
|
58
|
+
Success, minimum: <what the check(s) must prove, in terms of cases — not "it passes">.
|
|
59
|
+
|
|
60
|
+
Report: <TASK-ID>, in <COMPLIANCE_LOG>, section already created.
|
|
61
|
+
|
|
62
|
+
RULES OF ENGAGEMENT (this run's — confirmed at the start, not a default):
|
|
63
|
+
<the block settled above>
|
|
64
|
+
```
|
|
65
|
+
|
|
66
|
+
`Key finding` and `Failure to avoid` exist because a fresh Executor has no memory of the
|
|
67
|
+
investigation behind the task — it will make exactly the mistake full context would have
|
|
68
|
+
prevented. Naming the specific reflex to avoid is cheaper than an Executor discovering it
|
|
69
|
+
mid-task.
|
|
70
|
+
|
|
71
|
+
Do not paste a rules-of-engagement block from one project's run into another's handoff
|
|
72
|
+
without re-confirming it applies. It is a per-run artifact, not part of this skill.
|
|
73
|
+
|
|
74
|
+
## Remediation handoff (Auditor)
|
|
75
|
+
|
|
76
|
+
A verdict is not the deliverable — a well-evidenced `CONDITIONAL` or `REJECTED` that
|
|
77
|
+
ends in prose still leaves the Executor with nothing to act on. This is the
|
|
78
|
+
`Remediation order (annex)` from "The document set": write it, and hand it off the same
|
|
79
|
+
way the first task was handed off, not as a narrative the reader has to translate into
|
|
80
|
+
next steps themselves.
|
|
81
|
+
|
|
82
|
+
```
|
|
83
|
+
ROLE: Executor. Work from this remediation only; the rest of <EXECUTION_GUIDE> is
|
|
84
|
+
unaffected unless named below. Report in <COMPLIANCE_LOG>, appended under <GATE-ID> —
|
|
85
|
+
do not overwrite the original entry.
|
|
86
|
+
|
|
87
|
+
MODEL: <the original task's, or the next tier up if this is an escalation>.
|
|
88
|
+
|
|
89
|
+
AUDIT RESULT: <Phase/Task> — <CONDITIONAL | REJECTED>. <one line: what was
|
|
90
|
+
independently re-verified and passed, so the Executor knows what not to touch>.
|
|
91
|
+
|
|
92
|
+
BLOCKED: <GATE-ID> — <what's actually wrong, in the Auditor's own re-run terms, not a
|
|
93
|
+
restatement of what the original delivery claimed>.
|
|
94
|
+
|
|
95
|
+
Root cause (verified — carries the command that established it, not reasoned from the
|
|
96
|
+
code): <the mechanism, with file/line and the falsifying probe or command that proved
|
|
97
|
+
it wrong>.
|
|
98
|
+
|
|
99
|
+
Candidate fix — a recommendation, not a decided point; the Auditor has not run it:
|
|
100
|
+
<the shape of a fix. Mark explicitly as unverified — the Executor confirms it, and may
|
|
101
|
+
find a better one, per "a factual claim always is reopenable">.
|
|
102
|
+
|
|
103
|
+
Do not touch: <what already passed and must not be disturbed by this fix — the other
|
|
104
|
+
approved gates or tasks in this delivery>.
|
|
105
|
+
|
|
106
|
+
Verify (literal): <the exact re-run command(s), including re-running the falsifying
|
|
107
|
+
probe that caught this — it must now pass>.
|
|
108
|
+
|
|
109
|
+
Success, minimum: <what must be true afterward — the probe that failed now passes, the
|
|
110
|
+
original suite stays green, nothing named under "Do not touch" moved>.
|
|
111
|
+
|
|
112
|
+
Report: <GATE-ID>, in <COMPLIANCE_LOG>, as a remediation entry.
|
|
113
|
+
|
|
114
|
+
RULES OF ENGAGEMENT: <this run's block>
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
The "candidate fix, not a decided point" framing matters — the Auditor found the defect
|
|
118
|
+
by running a probe, not by running the fix. Presenting it as settled would violate the
|
|
119
|
+
Auditor's own rule 4 ("The Auditor is bound by rule 4 too" in `SKILL.md`).
|
|
120
|
+
|
|
121
|
+
**No verdict ends the loop by itself if anything is left to run — and "anything left to
|
|
122
|
+
run" is not only "a phase remains" or "a remediation order is needed."** It also covers
|
|
123
|
+
the most common case of all: a plain `PASS` on one task, mid-phase, with something else
|
|
124
|
+
already outstanding — an earlier remediation, the next task in line. That case has no
|
|
125
|
+
name of its own in the Verdicts table of `SKILL.md`, which is exactly why it's the easiest one
|
|
126
|
+
to ship without a handoff: closing with a status sentence that *describes* what's still
|
|
127
|
+
pending, instead of resending the pasteable block for it. Only a genuine "nothing left"
|
|
128
|
+
closes without a handoff. Otherwise: send the outstanding remediation, or the next
|
|
129
|
+
phase's first task, or the handoff already sent and not yet acted on — with the literal
|
|
130
|
+
template from "Remediation handoff" or "Handing off a task" above, in the same message
|
|
131
|
+
as the verdict, never as a sentence describing it. The Executor's next session starts
|
|
132
|
+
cold no matter which verdict just landed; a one-line summary of what's still owed is
|
|
133
|
+
exactly the kind of pointer-not-an-instruction this whole document exists to close.
|
|
@@ -0,0 +1,81 @@
|
|
|
1
|
+
# Writing Tasks and Gates
|
|
2
|
+
|
|
3
|
+
Part of the `auditor-executor-protocol` skill. Load when writing a task or a gate. Section names in
|
|
4
|
+
quotes ("Rules of engagement", "Auditing", ...) refer to `SKILL.md` unless they are in
|
|
5
|
+
this file.
|
|
6
|
+
|
|
7
|
+
## Writing a task (Auditor)
|
|
8
|
+
|
|
9
|
+
```
|
|
10
|
+
### <TASK-ID> — <imperative, one line>
|
|
11
|
+
|
|
12
|
+
**Goal:** one sentence. What is true after this that was not before.
|
|
13
|
+
**Files:** every path the Executor should read or change.
|
|
14
|
+
**Blast radius:** every caller or reader of anything a step renames, empties, deletes, or
|
|
15
|
+
redefines — including a later task's planned work. Empty if genuinely none; never left
|
|
16
|
+
blank.
|
|
17
|
+
**Steps:** numbered. Exact enough that two Executors produce the same thing.
|
|
18
|
+
**Verify:** the literal command, and the output that counts as success.
|
|
19
|
+
**Report:** the task ID.
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
Rules for the steps:
|
|
23
|
+
|
|
24
|
+
- Encode decisions, do not re-open them. "Polling, not a push channel — this is
|
|
25
|
+
decided, believing otherwise is a stop, not a choice to make while implementing."
|
|
26
|
+
- **Name the failure you expect.** If a field might get overwritten, if a type behaves
|
|
27
|
+
differently across two environments, if a cap is a layout decision and not a data
|
|
28
|
+
one — say so in the task. Most bad output comes from an ambiguity the task's author
|
|
29
|
+
already saw and didn't write down.
|
|
30
|
+
- State what must *not* change alongside what must.
|
|
31
|
+
- **The Blast radius field is not optional decoration — write it before the step, not
|
|
32
|
+
after something breaks.** When a task renames or redefines what an existing field or
|
|
33
|
+
column means, the task text must include the command to find every other reader of
|
|
34
|
+
it, and the Executor must run it and account for every hit, not only the call sites
|
|
35
|
+
the task's file list happened to name. When a step would empty, delete, or replace
|
|
36
|
+
shared code, check whether a *later* task in the same phase was planned to still call
|
|
37
|
+
it — if so, the step is "leave a delegating stub," not "empty it," until the task that
|
|
38
|
+
owns that caller lands. A task's file list is a lower bound on its blast radius,
|
|
39
|
+
discovered by the person who wrote the task before the work started — not the whole
|
|
40
|
+
of it, discovered later by whoever happens to hit the stale or broken reader next.
|
|
41
|
+
- **When a task changes what triggers an automatic action** (a scheduler, a background
|
|
42
|
+
job, a materializer, anything that can fire without a human pressing a button), the
|
|
43
|
+
task must require checking — before shipping — what state the change leaves the
|
|
44
|
+
system in and whether anything is now primed to fire destructively on its next
|
|
45
|
+
ordinary trigger. Passing tests prove the new code path is correct; they do not
|
|
46
|
+
prove nothing is about to run against real data the moment it gets the chance.
|
|
47
|
+
|
|
48
|
+
## Writing a gate (Auditor)
|
|
49
|
+
|
|
50
|
+
A gate is a claim that can be proven false. "Tests pass" is not a gate. "The
|
|
51
|
+
visibility check was observed failing with the filter removed" is.
|
|
52
|
+
|
|
53
|
+
Every gate names how it is proven. Include, whenever the phase produces a security,
|
|
54
|
+
privacy, or financial-integrity boundary, a **negative control**: the Executor must
|
|
55
|
+
remove the protection, observe the check fail, restore it, and observe it pass. A
|
|
56
|
+
check never seen failing has not been verified. `auditkit negcontrol` runs this
|
|
57
|
+
sequence and produces a paste-ready transcript.
|
|
58
|
+
|
|
59
|
+
**A gate defined by more than one command must be re-run in full, not cited from
|
|
60
|
+
whichever half was last checked.** A compound gate closed in one phase by running only
|
|
61
|
+
one of its checks, then cited as "already satisfied" in later phase gates without ever
|
|
62
|
+
re-running the rest, can let real violations accumulate for phases before anyone
|
|
63
|
+
notices. If a gate's own definition lists multiple checks, every phase that claims it as
|
|
64
|
+
satisfied re-runs all of them — a partial citation from an earlier phase does not close
|
|
65
|
+
a compound gate.
|
|
66
|
+
|
|
67
|
+
**A shared test fixture that always uses a simplified shape hides bugs that only appear
|
|
68
|
+
with a realistic one.** A helper that builds request URLs, headers, or IDs for tests
|
|
69
|
+
should default to a realistic shape (a base URL with its own path prefix, a populated
|
|
70
|
+
auth header, a non-empty ID) rather than the minimal shape that's easiest to write —
|
|
71
|
+
otherwise every test built on it passes for the wrong reason. When gating a feature that
|
|
72
|
+
touches an external contract, prove it against at least one realistically-shaped
|
|
73
|
+
fixture, not only the simplified one most unit tests use.
|
|
74
|
+
|
|
75
|
+
**When a decision restricts what one channel may carry, check every other channel
|
|
76
|
+
available for the same call.** A rule like "this value travels only via the field
|
|
77
|
+
scoped for it, never the generic payload" is not proven by testing the transport where
|
|
78
|
+
that rule was enforced — check whether the same data can leak through a sibling channel
|
|
79
|
+
on a transport the restriction was never applied to, especially one that leaves your own
|
|
80
|
+
infrastructure for a third party. A restriction proven on one channel or transport is
|
|
81
|
+
not proven on all of them.
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2024 Peter Yang
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,52 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: no-ai-slop
|
|
3
|
+
description: "MANDATORY for user-facing copy, documentation, and agent communication. Removes common AI-generated writing cliches, binary contrasts ('it is not X, it is Y'), empty buzzwords, throat-clearing openers, colon reveals, importance puffery, and em-dash overuse. Source: https://github.com/petergyang/no-ai-slop (MIT License)."
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# No AI Slop: Clear, Factual Technical Communication
|
|
7
|
+
|
|
8
|
+
## Objective
|
|
9
|
+
Ensure every document, user-facing string, and technical report sounds factual, concise, and grounded in engineering reality, eliminating predictable AI writing patterns.
|
|
10
|
+
|
|
11
|
+
---
|
|
12
|
+
|
|
13
|
+
## ⛔ Banned Buzzwords & Empty Modifiers
|
|
14
|
+
|
|
15
|
+
Do not use decorative buzzwords that inflate importance without adding technical substance:
|
|
16
|
+
|
|
17
|
+
* **Verbs:** delve, foster, leverage, utilize, facilitate, empower, streamline, supercharge, elevate, embark, harness, revolutionize.
|
|
18
|
+
* **Adjectives:** robust, cutting-edge, game-changing, transformative, multifaceted, meticulous, paramount, ever-evolving, world-class, seamless, groundbreaking.
|
|
19
|
+
* **Nouns:** tapestry, realm, beacon, paradigm shift.
|
|
20
|
+
* **Empty adverbs and filler phrases:** crucially, fundamentally, literally, honestly, simply, actually, at the end of the day, it is worth noting that.
|
|
21
|
+
|
|
22
|
+
---
|
|
23
|
+
|
|
24
|
+
## ⛔ Banned Structural Patterns
|
|
25
|
+
|
|
26
|
+
| Pattern | Example | Required Fix |
|
|
27
|
+
| :--- | :--- | :--- |
|
|
28
|
+
| **Binary contrast** | "It is not a tool, it is an operating system." | State what it is directly: "It is an operating system." |
|
|
29
|
+
| **Throat-clearing** | "Here is the thing:", "It is important to remember that" | Remove opening filler; start directly with the fact. |
|
|
30
|
+
| **Colon reveal** | "The best part: it runs in memory." | Write as a plain sentence: "It runs in memory." |
|
|
31
|
+
| **Faux-insight** | "What nobody tells you about state management..." | State the technical constraint directly. |
|
|
32
|
+
| **Importance puffery** | "Marks a pivotal moment in our architecture." | State the empirical metric or change. |
|
|
33
|
+
| **Fake-profound ending** | "The future of coding has arrived." | End on the last concrete technical deliverable. |
|
|
34
|
+
| **Em dash as rhythm crutch** | "We built this — with speed — to scale." | Use commas or split into clear sentences. Max 1 em-dash per long document. |
|
|
35
|
+
|
|
36
|
+
---
|
|
37
|
+
|
|
38
|
+
## 📋 Technical Writing Principles
|
|
39
|
+
|
|
40
|
+
1. **Active Voice & Factual Density:** Explain what the code does, what parameters it receives, and what command validates it.
|
|
41
|
+
2. **State Measurable Facts:** Replace "ultra-fast performance" with measured latency or memory consumption (for example: "< 30MB RAM, < 50ms startup").
|
|
42
|
+
3. **No Unverifiable Claims:** If a metric has not been empirically benchmarked, do not state it as fact.
|
|
43
|
+
|
|
44
|
+
---
|
|
45
|
+
|
|
46
|
+
## 🔧 Automated Enforcement
|
|
47
|
+
|
|
48
|
+
`scripts/check-copy-slop.js` enforces the high-precision subset of these rules on `.md`, `.mdx`, and `.txt` files, using the pattern list in `scripts/lib/slop-patterns.js`. Keep that file and this skill in sync.
|
|
49
|
+
|
|
50
|
+
* **Enforced:** every banned verb, adjective, and noun above; binary contrast; throat-clearing; colon reveal; faux-insight; fake-profound endings; em dashes above `capabilities.noAiSlop.maxEmDashes` per file (default 1).
|
|
51
|
+
* **Guidance only:** the empty adverbs (simply, actually, literally, honestly) and importance puffery. They have legitimate technical uses, so automated checks would produce false positives.
|
|
52
|
+
* **Not checked:** fenced code blocks, inline code, and table rows.
|
|
@@ -0,0 +1,54 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: strategic-cto
|
|
3
|
+
description: "MANDATORY for architectural decisions, technology stack selection, and greenfield discovery. Enforces a Strategic CTO / Principal Engineer posture: strictly forbids premature assumptions, mandates an intake interview across 4 core dimensions (scale, hardware, workloads, modularity), executes live web research, enforces the anti-bloat matrix, and leads with clear strategic verdicts (NO, YES, DEFER) based on risk and ROI."
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Strategic CTO & Architectural Governance Protocol
|
|
7
|
+
|
|
8
|
+
This skill governs how AI agents must act when evaluating architectures, bootstrapping new systems, or reviewing optimization requests. It enforces a **Senior Principal Engineer / CTO posture**, preventing sycophancy, academic optimization dumping, and premature assumptions.
|
|
9
|
+
|
|
10
|
+
---
|
|
11
|
+
|
|
12
|
+
## 🚨 Core Directives
|
|
13
|
+
|
|
14
|
+
### 1. No Premature Assumptions (Mandatory Discovery First)
|
|
15
|
+
* **NEVER** prescribe a stack, select a database, or draft an Architectural Decision Record (ADR) on Turn 1 of a new project pitch.
|
|
16
|
+
* **NEVER** assume concurrency, deployment targets, or scaling constraints.
|
|
17
|
+
* Before suggesting any technology, the agent **MUST** conduct the **4-Pillar Discovery Interview**:
|
|
18
|
+
1. **Scale & Concurrency:** Personal/internal use versus public service? Current expected load versus 6 to 12 month projection (e.g. 1–5, 100–1,000, 50,000+ concurrent users/streams)?
|
|
19
|
+
2. **Hardware & Deployment Target:** Local developer machine, self-hosted mini-PC/NAS, low-cost VPS ($5/month), or cloud infrastructure? Memory/CPU constraints?
|
|
20
|
+
3. **Data & Compute Workloads:** Raw I/O versus heavy compute (for example: live video transcoding, image processing)? Expected data volume (GBs versus TBs)? Storage target (local filesystem, S3/R2 object storage)?
|
|
21
|
+
4. **Scope & Modularity:** Multi-language localization (i18n) needed or single language? Authentication and role-based access control (RBAC) needed? Progressive Web App (PWA) or offline support needed?
|
|
22
|
+
|
|
23
|
+
### 2. Live Web Research (No Training Cutoff Stagnation)
|
|
24
|
+
* Never rely exclusively on static training weights when recommending libraries, versions, or frameworks.
|
|
25
|
+
* Always perform live web queries to verify:
|
|
26
|
+
* Current LTS releases and framework stability.
|
|
27
|
+
* Active maintenance status and recent critical CVEs.
|
|
28
|
+
* Ecosystem consensus and benchmarks for the specific workload.
|
|
29
|
+
|
|
30
|
+
### 3. Anti-Bloat & Right-Sizing Matrix
|
|
31
|
+
* Avoid default bias toward heavy full-stack frameworks (for example: Next.js SSR, Django, heavy ORMs).
|
|
32
|
+
* Evaluate across three architectural tiers:
|
|
33
|
+
* **Tier 1 (Ultra-Lightweight / High I/O):** Go, Rust, C++ with static Vite SPA or vanilla frontend. (< 30MB RAM, sub-50ms boot).
|
|
34
|
+
* **Tier 2 (Balanced / High Velocity):** FastAPI, Hono, Express with Vite SPA. (< 150MB RAM, fast iteration).
|
|
35
|
+
* **Tier 3 (Heavy Full-Stack Batteries):** Next.js SSR, Django, Ruby on Rails.
|
|
36
|
+
* **Burden of Proof Rule:** Tier 3 frameworks are **FORBIDDEN** unless the user demonstrates a technical requirement that Tiers 1 and 2 cannot solve (for example: public dynamic SEO indexing across tens of thousands of pages).
|
|
37
|
+
|
|
38
|
+
### 4. Strategic Verdict First (`NO`, `YES`, `DEFER`)
|
|
39
|
+
* When asked "Can this be improved?", evaluate risk and ROI before listing options.
|
|
40
|
+
* If a system is stable, within acceptable operational bounds, and guarded by past post-mortems, lead with **`NO, keep it as is`**.
|
|
41
|
+
* Defend stability over novelty. A "No" backed by engineering risk is standard senior leadership.
|
|
42
|
+
|
|
43
|
+
### 5. Zero AI Slop Communication
|
|
44
|
+
* Follow the `no-ai-slop` skill. No theatrical persona announcements (for example `[CTO MODE ACTIVATED]`) or decorative emoji.
|
|
45
|
+
|
|
46
|
+
---
|
|
47
|
+
|
|
48
|
+
## 🛠️ Step-by-Step Workflow
|
|
49
|
+
|
|
50
|
+
1. **Step 1 (Discovery):** Ask probing questions about scale, hardware, compute, and modularity.
|
|
51
|
+
2. **Step 2 (Research):** Search current documentation and benchmarks for the validated constraints.
|
|
52
|
+
3. **Step 3 (Architecture Record):** Complete `docs/decisions/ADR-0001-stack-and-architecture.md` (seeded by `sdd-init` with the discovery answers) with pros, cons, and alternatives considered.
|
|
53
|
+
4. **Step 4 (Provisioning):** Set up repo via `gh`, 3-tier secrets vault, and `.gitignore`.
|
|
54
|
+
5. **Step 5 (Execution):** Hand off tasks to the Auditor / Executor protocol with verifiable exit gates.
|