agentic-sdd-framework 1.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (57) hide show
  1. package/.agents/AGENTS.template.md +59 -0
  2. package/.agents/CONTEXT.template.md +41 -0
  3. package/.agents/ENTRYPOINT.template.md +31 -0
  4. package/.agents/skills/ast-navigator/SKILL.md +31 -0
  5. package/.agents/skills/ast-navigator/adapters/ast-grep.md +18 -0
  6. package/.agents/skills/ast-navigator/adapters/graphify.md +19 -0
  7. package/.agents/skills/ast-navigator/adapters/lsp.md +16 -0
  8. package/.agents/skills/ast-navigator/adapters/ripgrep.md +19 -0
  9. package/.agents/skills/auditor-executor-protocol/SKILL.md +410 -0
  10. package/.agents/skills/auditor-executor-protocol/references/autonomous-mode.md +144 -0
  11. package/.agents/skills/auditor-executor-protocol/references/failure-modes-and-example.md +103 -0
  12. package/.agents/skills/auditor-executor-protocol/references/handoffs.md +133 -0
  13. package/.agents/skills/auditor-executor-protocol/references/tasks-and-gates.md +81 -0
  14. package/.agents/skills/no-ai-slop/LICENSE +21 -0
  15. package/.agents/skills/no-ai-slop/SKILL.md +52 -0
  16. package/.agents/skills/strategic-cto/SKILL.md +54 -0
  17. package/CHANGELOG.md +117 -0
  18. package/LICENSE +21 -0
  19. package/README.md +244 -0
  20. package/docs/SPEC_TEMPLATE.md +78 -0
  21. package/docs/decisions/ADR_TEMPLATE.md +49 -0
  22. package/docs/guidelines/AST_NAVIGATION.md +51 -0
  23. package/docs/guides/AGENT_CREDENTIALS.md +75 -0
  24. package/docs/guides/GITHUB_CLI_SETUP.md +74 -0
  25. package/docs/incidents/0000-00-00-incident-template.md +35 -0
  26. package/docs/roadmap/templates/compliance-log.md +37 -0
  27. package/docs/roadmap/templates/execution-guide.md +75 -0
  28. package/docs/roadmap/templates/plan-of-record.md +49 -0
  29. package/package.json +49 -0
  30. package/scripts/check-copy-slop.js +120 -0
  31. package/scripts/check-file-size.js +66 -0
  32. package/scripts/check-spec.js +201 -0
  33. package/scripts/check-system-prerequisites.js +133 -0
  34. package/scripts/check-versions.js +50 -0
  35. package/scripts/dev/fuzz-spec-markup.js +123 -0
  36. package/scripts/dev/set-npm-publish-token.sh +40 -0
  37. package/scripts/dev/sync-vendored.js +94 -0
  38. package/scripts/install-git-hooks.js +103 -0
  39. package/scripts/lib/cli.js +60 -0
  40. package/scripts/lib/config.js +111 -0
  41. package/scripts/lib/git.js +211 -0
  42. package/scripts/lib/markdown.js +46 -0
  43. package/scripts/lib/provision.js +323 -0
  44. package/scripts/lib/runner.js +70 -0
  45. package/scripts/lib/sdd.config.schema.json +213 -0
  46. package/scripts/lib/slop-patterns.js +57 -0
  47. package/scripts/lib/spec-markup.js +346 -0
  48. package/scripts/lib/spec.js +226 -0
  49. package/scripts/lib/state.js +107 -0
  50. package/scripts/lib/vendor/README.md +11 -0
  51. package/scripts/lib/vendor/markdown-it.LICENSE +22 -0
  52. package/scripts/lib/vendor/markdown-it.min.js +3 -0
  53. package/scripts/quality-gate.js +151 -0
  54. package/scripts/sdd-init.js +245 -0
  55. package/scripts/sdd-verify.js +176 -0
  56. package/scripts/verify-no-secrets.js +216 -0
  57. package/sdd.config.json +33 -0
@@ -0,0 +1,59 @@
1
+ # Agent Constitution and Behavioral Guidelines
2
+
3
+ This document establishes the non-negotiable operating rules for AI coding agents working in this repository. Every agent must read and adhere to these directives before executing tasks.
4
+
5
+ ---
6
+
7
+ ## 1. Discovery First (No Premature Assumptions)
8
+ * **Rule:** Before recommending architectures, selecting frameworks, or generating code on a new initiative, the agent must execute the 4-Pillar Discovery Interview (Scale/Concurrency, Hardware/Deployment, Workload/Compute, Modularity).
9
+ * **Why this rule exists:**
10
+ > [Document the incident or rationale here. For example: An agent previously recommended a heavy full-stack framework for a lightweight streaming utility, introducing unnecessary runtime bloat.]
11
+
12
+ ---
13
+
14
+ ## 2. Evidence-Driven Debugging & Diagnostics
15
+ * **Rule:** Never guess root causes or apply speculative patches. Inspect log files, inspect command output, and run diagnostics before altering code.
16
+ * **Why this rule exists:**
17
+ > [Document the incident or rationale here. For example: Speculative patches often introduce regression cascades that obscure the original defect.]
18
+
19
+ ---
20
+
21
+ ## 3. Mandatory Verification Before Certification
22
+ * **Rule:** A task or phase is not complete until its explicit verification command exits with code 0. Reading code visually is never a substitute for running the code.
23
+ * **Why this rule exists:**
24
+ > [Document the incident or rationale here. For example: Clean syntax does not guarantee functional runtime behavior or integration success.]
25
+
26
+ ---
27
+
28
+ ## 4. Closed-Network Testing Isolation
29
+ * **Rule:** Automated test suites must never contact external internet hosts. All external integrations must be mocked or gated on environment variables. Loopback testing is permitted for local servers.
30
+ * **Why this rule exists:**
31
+ > [Document the incident or rationale here. For example: Test suites that hit third-party APIs can cause hard bounces, rate limits, and unexpected billing exposure.]
32
+
33
+ ---
34
+
35
+ ## 5. Zero-Trust Secrets Management
36
+ * **Rule:** Agents must never request API keys or credentials in chat prompts. Secrets must be read directly from the Tier 3 Vault (`~/secrets/<app>/.vault`) or environment variables. Never commit secrets to Git. Tier model: `docs/guides/AGENT_CREDENTIALS.md`.
37
+ * **Why this rule exists:**
38
+ > [Document the incident or rationale here. For example: Chat logs and git histories are permanently recorded; leaking credentials requires immediate key revocation.]
39
+
40
+ ---
41
+
42
+ ## 6. Scope Bounding & Atomic Progression
43
+ * **Rule:** Execute one task at a time in strict sequence. Do not refactor unrelated files or perform out-of-scope cleanups without explicit Auditor authorization.
44
+ * **Why this rule exists:**
45
+ > [Document the incident or rationale here. For example: Unbounded edits make regression tracing difficult and invalidate concurrent session handoffs.]
46
+
47
+ ---
48
+
49
+ ## 7. Factual Technical Copy (No AI Slop)
50
+ * **Rule:** All documentation, user-facing copy, and commit messages must be factual, concise, and dense, following the `no-ai-slop` skill. The quality gate enforces its banned patterns.
51
+ * **Why this rule exists:**
52
+ > [Document the incident or rationale here. For example: Promotional puffery and marketing contrasts obscure technical reality and degrade documentation clarity.]
53
+
54
+ ---
55
+
56
+ ## 8. Author Attribution & Integrity
57
+ * **Rule:** Never attach `Co-Authored-By:` trailers crediting AI assistants to Git commits. The repository owner is the sole author.
58
+ * **Why this rule exists:**
59
+ > [Document the incident or rationale here. For example: Project contributors retain 100% ownership and commit graph clarity without AI tooling noise.]
@@ -0,0 +1,41 @@
1
+ # Operational Context and Memory Registry
2
+
3
+ This file records the cumulative architectural decisions, production incidents, and operational learnings for this repository. AI agents read this file to avoid repeating past mistakes.
4
+
5
+ ---
6
+
7
+ ## 1. Project Summary & Active Capabilities
8
+
9
+ * **Project Name:** [Project Name]
10
+ * **Specification Mode:** Lite | Rigor
11
+ * **Primary Runtime:** [e.g. Node.js 24 LTS, Go 1.23, Python 3.12]
12
+ * **Architecture Style:** Clean Architecture | Standard MVC | Modular Monolith
13
+ * **Active Capabilities:**
14
+ * No-AI-Slop Linter: [Enabled / Disabled]
15
+ * AST Navigator Adapter: [graphify / ast-grep / ripgrep / lsp]
16
+ * Internationalization (i18n): [Enabled / Disabled]
17
+ * Progressive Web App (PWA): [Enabled / Disabled]
18
+
19
+ ---
20
+
21
+ ## 2. Incident & Root-Cause Registry
22
+
23
+ Whenever a production regression, test flake, or operational issue is resolved, record the incident below. Every incident must produce a verifiable rule in `.agents/AGENTS.md`.
24
+
25
+ ### Template:
26
+ ```markdown
27
+ ### [YYYY-MM-DD] Incident: [Short Title]
28
+ * **Symptom:** What broke and where was it observed.
29
+ * **Root Cause:** The underlying mechanical or architectural flaw.
30
+ * **Remediation:** The code change or configuration fix applied.
31
+ * **Inviolable Rule Created:** Link to the corresponding rule added to AGENTS.md.
32
+ ```
33
+
34
+ ---
35
+
36
+ ## 3. Active Technical Debt & Non-Goals
37
+
38
+ Document known constraints, deferred refactors, and explicit non-goals to prevent agents from attempting unsolicited cleanups.
39
+
40
+ * [Non-Goal 1]: [Description]
41
+ * [Technical Debt 1]: [Description and tracking issue]
@@ -0,0 +1,31 @@
1
+ <!-- sdd:managed. Generated by sdd-init; rerun it to refresh. Put project rules in .agents/AGENTS.md. -->
2
+ # Agent Entry Point: {{PROJECT_NAME}}
3
+
4
+ This project follows the Agentic SDD Framework. Before any task, read:
5
+
6
+ 1. `.agents/AGENTS.md`: non-negotiable rules for this repository.
7
+ 2. `.agents/CONTEXT.md`: incident registry, technical debt, and non-goals.
8
+ 3. {{SPEC_LOCATION}}
9
+
10
+ ## Skills
11
+
12
+ Load a skill from `.agents/skills/<name>/SKILL.md` only when its trigger applies.
13
+
14
+ | Skill | Load when |
15
+ | :--- | :--- |
16
+ | `strategic-cto` | Choosing a stack or architecture, or answering "can this be improved?" |
17
+ | `auditor-executor-protocol` | Work spans several phases or sessions, or another agent implements the plan |
18
+ | `no-ai-slop` | Writing documentation, user-facing copy, or commit messages |
19
+ | `ast-navigator` | Exploring code. Active adapter: `{{AST_ADAPTER}}` (`.agents/skills/ast-navigator/adapters/{{AST_ADAPTER}}.md`) |
20
+
21
+ ## Project Facts
22
+
23
+ * Runtime: {{RUNTIME}}
24
+ * Specification mode: {{SPEC_MODE}}
25
+ * Configuration: `sdd.config.json`
26
+
27
+ ## Before Every Push
28
+
29
+ Run `{{GATE_COMMAND}}`. It must exit 0. The pre-push hook runs it automatically, and it checks the specification as well as the code: a checked task or `DONE` report without pasted evidence fails.
30
+
31
+ {{SPEC_CHECK}}
@@ -0,0 +1,31 @@
1
+ ---
2
+ name: ast-navigator
3
+ description: "MANDATORY for codebase exploration, symbol lookups, and dependency tracing. Enforces AST-first structural navigation before falling back to full-text grep, minimizing token usage and context clutter. Supports pluggable adapters: Graphify, ast-grep, ripgrep, and LSP/SCIP."
4
+ ---
5
+
6
+ # AST Navigator Protocol: Token-Efficient Code Exploration
7
+
8
+ ## Objective
9
+ Enable AI coding agents to navigate complex codebases by querying Abstract Syntax Tree (AST) representations and dependency graphs first, drastically reducing LLM context consumption compared to reading entire files or running unbounded text searches.
10
+
11
+ ---
12
+
13
+ ## 🧭 Navigation Hierarchy
14
+
15
+ 1. **Tier 1 (AST / Graph Query):** Query structural dependencies, class hierarchies, and symbol definitions using the configured adapter in `sdd.config.json`.
16
+ 2. **Tier 2 (Fallback to Grep):** Fall back to targeted regex grep only when:
17
+ * The AST adapter returns no matches.
18
+ * The graph index is known to be stale and cannot be updated immediately.
19
+ * Exact line-level character sequences (string literals, CSS rules, comments) are required.
20
+ 3. **Tier 3 (File Inspection):** Read specific line slices (`view_file`) only after locating the exact symbol and line numbers. Never dump entire 1,000-line files into context without target line bounds.
21
+
22
+ ---
23
+
24
+ ## 🔌 Active Adapter
25
+
26
+ Read `capabilities.astNavigation.adapter` in `sdd.config.json`, then load only `adapters/<adapter>.md` for its commands. Available adapters: `ast-grep`, `graphify`, `ripgrep`, `lsp`. The comparison used to choose one lives in `docs/guidelines/AST_NAVIGATION.md`.
27
+
28
+ ---
29
+
30
+ ## 🚨 Verification Rule Before Claiming Compliance
31
+ A missing edge or symbol in an AST index is **not** definitive proof that a dependency does not exist. Before certifying architectural compliance or absence of a vulnerability, always verify the actual import statements of the target file.
@@ -0,0 +1,18 @@
1
+ # Adapter: ast-grep
2
+
3
+ * **Upstream Creator:** Herrington Darkholme (`@HerringtonDarkholme`)
4
+ * **Repository:** [https://github.com/ast-grep/ast-grep](https://github.com/ast-grep/ast-grep)
5
+ * **License:** MIT License
6
+ * **Runtime Requirement:** Native binary (Rust-powered, installable via `npm i -g @ast-grep/cli` or `brew install ast-grep`)
7
+
8
+ ## Capabilities and Mechanics
9
+ ast-grep performs structural code searches and transformations using Tree-sitter syntax trees. Unlike plain regex, it matches syntax constructs regardless of formatting, indentation, or comments.
10
+
11
+ ## Common Operational Commands
12
+ * **Search AST Pattern:** `sg run -p '$FUNCTION($$$ARGS)' -l typescript`
13
+ * **Find Class Definitions:** `sg run -p 'class $NAME extends $BASE { $$$ }'`
14
+ * **Scan Rules:** `sg scan` (executes pre-configured YAML linting rules).
15
+
16
+ ## Best Practices
17
+ * Use for lightweight environments where Python dependencies are undesirable.
18
+ * Ideal for CI/CD structural rule enforcement and automated AST refactors.
@@ -0,0 +1,19 @@
1
+ # Adapter: Graphify
2
+
3
+ * **Upstream Organization:** Graphify Labs (`@Graphify-Labs`)
4
+ * **Repository:** [https://github.com/Graphify-Labs/graphify](https://github.com/Graphify-Labs/graphify)
5
+ * **License:** Apache License 2.0
6
+ * **Runtime Requirement:** Python 3.10+ (`pip install graphifyy`)
7
+
8
+ ## Capabilities and Mechanics
9
+ Graphify extracts relational knowledge graphs of symbols, imports, and calls across multi-language repositories, generating interactive visual graphs and structured JSON query outputs.
10
+
11
+ ## Common Operational Commands
12
+ * **Update Graph:** `graphify update .` (or configured project script `npm run graph:update`).
13
+ * **Check Index Freshness:** `graphify check` (validates graph commit against git HEAD).
14
+ * **Query Symbol Path:** `graphify path <source-node> <target-node>`.
15
+ * **Explain Subgraph:** `graphify explain <symbol-name>`.
16
+
17
+ ## Best Practices
18
+ * Always regenerate the graph after major refactors or file deletions.
19
+ * Exclude build noise directories (`dist/`, `build/`, `.venv/`) from the index.
@@ -0,0 +1,16 @@
1
+ # Adapter: Language Server Protocol (LSP) / SCIP
2
+
3
+ * **Upstream Organization:** SCIP Code (`@scip-code`), originally created by Sourcegraph
4
+ * **Repository:** [https://github.com/scip-code/scip](https://github.com/scip-code/scip)
5
+ * **License:** Apache License 2.0
6
+ * **Runtime Requirement:** Language-specific indexers (`scip-typescript`, `scip-go`, etc.)
7
+
8
+ ## Capabilities and Mechanics
9
+ SCIP (Source Code Intelligence Protocol) indexes definitions, references, and hover types across strongly typed codebases. It enables exact compiler-level cross-file navigation without heuristics.
10
+
11
+ ## Common Operational Commands
12
+ * **Index Project:** `scip-typescript index`
13
+ * **Query References:** Query through SCIP CLI or LSP client interfaces.
14
+
15
+ ## Best Practices
16
+ * Use in large enterprise repositories where type relationships and cross-package references must be verified with compiler precision.
@@ -0,0 +1,19 @@
1
+ # Adapter: ripgrep
2
+
3
+ * **Upstream Creator:** Andrew Gallant (`@BurntSushi`)
4
+ * **Repository:** [https://github.com/BurntSushi/ripgrep](https://github.com/BurntSushi/ripgrep)
5
+ * **License:** MIT License / UNLICENSE
6
+ * **Runtime Requirement:** Native binary (`rg`)
7
+
8
+ ## Capabilities and Mechanics
9
+ ripgrep is an ultra-fast line-oriented search tool that respects `.gitignore` rules by default. It provides the universal baseline fallback when AST parsers cannot resolve a query or when searching non-code files (markdown, JSON, YAML).
10
+
11
+ ## Common Operational Commands
12
+ * **Find Literal String:** `rg -F "exact_symbol_name"`
13
+ * **Case-Insensitive Search:** `rg -i "pattern"`
14
+ * **Filter by File Extension:** `rg "pattern" -g "*.ts"`
15
+ * **List Matching Files Only:** `rg -l "pattern"`
16
+
17
+ ## Best Practices
18
+ * Always use `--type` or `-g` globs to restrict search scope when querying large repositories.
19
+ * Avoid dumping thousands of lines into agent context; pair with `head` or line count limits.
@@ -0,0 +1,410 @@
1
+ ---
2
+ name: auditor-executor-protocol
3
+ description: "Runs multi-phase work as a two-role protocol: an Auditor who writes numbered, verifiable tasks and signs off phases, and an Executor who implements one task at a time and reports evidence. Use whenever a piece of work is large enough to span several phases or sessions, whenever handing a plan to another agent (subagent, spawned session, or a different tool) to implement, or when asked to audit/verify work another agent reports as complete. Also use when a plan needs to become executable instructions rather than a discussion document. Covers: the document set, rules of engagement, task and gate format, how to audit by re-running rather than reading, negative controls, verdict vocabulary, the annex pattern for correcting an order that is already in flight, and an optional operating mode chosen per run — Guided (the owner relays each handoff) or Autonomous (the Auditor launches Executors as subagents with an explicit per-task model tier, answers their BLOCKED questions itself, and interrupts the owner only for a numbered critical-risk list). Triggers: 'audit', 'auditor', 'executor', 'phased plan', 'exit criteria', 'gate', 'compliance log', 'remediation', 'handoff', 'have another agent implement this', 'verify what was reported', 'sign off', 'autonomous mode', 'guided mode'. NOT for reviewing a diff or a pull request for bugs — that's a code review; this protocol audits an execution report against pre-declared gates, and governs how multi-phase work is handed off and signed off."
4
+ ---
5
+
6
+ # Auditor / Executor Protocol
7
+
8
+ Two roles, one paper trail. The Auditor decides what "done" means and proves it
9
+ independently. The Executor implements and reports evidence. Neither does the other's job.
10
+
11
+ This exists because a plan handed to an implementing agent is not instructions until
12
+ someone makes it unambiguous, and a report that says "all green" is not verification
13
+ until someone re-runs it. Both failures are common and both are expensive once they
14
+ compound across dozens of tasks.
15
+
16
+ A companion CLI, `auditkit`, operationalizes the mechanical parts of this: scaffolding
17
+ the document set, checking the task list and the log agree with each other, running a
18
+ negative control end to end, and printing where a run stands. Use it if it's installed;
19
+ follow this document either way.
20
+
21
+ ## When to use this
22
+
23
+ - Work spanning more than one phase or more than one session.
24
+ - Handing a plan to another agent to implement — a subagent, a spawned session, a
25
+ different tool, another person.
26
+ - Auditing work someone else reports as complete.
27
+ - Turning a roadmap into something executable.
28
+
29
+ Not for single-session tasks you are doing yourself. The overhead only pays off when the
30
+ implementer is not the person who wrote the plan.
31
+
32
+ ## The two roles
33
+
34
+ | | Auditor | Executor |
35
+ |---|---|---|
36
+ | Owns | The instructions, the gates, the verdicts | The implementation, the evidence |
37
+ | Never | Writes feature code | Redesigns, or decides scope |
38
+ | Output | Task expansions, verdicts, remediation orders | Working code, pasted command output |
39
+
40
+ Whoever writes the plan is the Auditor. Say which role you are holding at the start of a
41
+ session so it does not drift.
42
+
43
+ ## Reference files (load on demand)
44
+
45
+ The core rules are in this file. Load a reference file only when you reach the step it
46
+ covers; each one is self-contained and refers back to sections of this file by name.
47
+
48
+ | File | Load when |
49
+ |---|---|
50
+ | [`references/tasks-and-gates.md`](references/tasks-and-gates.md) | Writing a task or a gate (Auditor) |
51
+ | [`references/handoffs.md`](references/handoffs.md) | Handing a task to an Executor, or writing a remediation after a verdict that is not a clean `APPROVED` (Auditor) |
52
+ | [`references/autonomous-mode.md`](references/autonomous-mode.md) | The run is in Autonomous mode, or you are writing the Auditor brief that starts one |
53
+ | [`references/failure-modes-and-example.md`](references/failure-modes-and-example.md) | Auditing a delivery, or you want to see a filled-in run end to end |
54
+
55
+ ## Operating mode: Guided or Autonomous (chosen at the start of the run)
56
+
57
+ The mode is optional and the owner's choice. Two ways to run the protocol:
58
+
59
+ | | Guided mode | Autonomous mode |
60
+ |---|---|---|
61
+ | Who launches Executors | The owner pastes each handoff into a fresh session | The Auditor, as subagents |
62
+ | Who answers `BLOCKED` | The owner, or the Auditor through the owner | The Auditor, with a decision entry in the log |
63
+ | Owner's role | Relays every message, sees every verdict | Receives one short report per closed phase; interrupted only for the critical-risk list |
64
+ | Fits when | The owner wants to watch and approve each step, the work is new territory, or the Auditor has no subagent tool | The plan is settled and the owner wants it run end to end |
65
+
66
+ How the mode is chosen:
67
+
68
+ - **If the owner already named it** ("autonomous", "guided", or an unambiguous
69
+ equivalent like "don't interrupt me for every step"), use it. Do not ask again.
70
+ - **If not, ask once, before task 1**, as a single two-option question with a one-line
71
+ recommendation (Guided for a first run in unfamiliar code or anything touching money
72
+ or production data with no rehearsal; Autonomous for a settled plan with strong gates).
73
+ This is one of the few questions that is genuinely the owner's: it decides how often
74
+ they get interrupted for the rest of the run.
75
+ - **If the environment has no subagent tool**, Autonomous is not available; say so and
76
+ run Guided.
77
+ - Record the choice as a governance decision in the plan of record, next to the rules of
78
+ engagement. Switching modes mid-run is allowed only when the owner asks, and is
79
+ recorded by annex. The Auditor never switches itself from Guided to Autonomous; it may
80
+ drop from Autonomous to Guided only through an item on the critical-risk list.
81
+
82
+ Everything else in this document applies to both modes. Autonomous mode adds the rules in
83
+ `references/autonomous-mode.md`; it removes none. Re-running, negative controls and separate
84
+ process/code findings matter more when nobody is watching, not less.
85
+
86
+ ## The document set
87
+
88
+ Four documents. Keep them separate; merging them is how the instructions turn back into
89
+ a discussion.
90
+
91
+ 1. **Plan of record** — the *what* and the *why*. Phases, decisions, trade-offs. Nobody
92
+ implements from this. Carries the **deferred items ledger** (below).
93
+ 2. **Execution guide** — the *how*. Numbered tasks (`P<phase>-T<n>`, or any scheme with a
94
+ stable, greppable ID), each with files, steps, a verification command, and its
95
+ expected output. Gates (`P<phase>-G<n>`) close each phase. Carries the
96
+ **reserved-to-Auditor steps** list (below).
97
+ 3. **Compliance log** — where the Executor reports. Pre-generate one empty row per task
98
+ and gate ID so nothing can be quietly skipped.
99
+ 4. **Remediation order (annex)** — written by the Auditor after an audit that isn't a
100
+ clean `APPROVED`. Self-contained: the Executor must not need the audit conversation
101
+ or the full log to act on it. Template and reasoning: "Remediation handoff" in `references/handoffs.md`.
102
+
103
+ **Annexes** supersede a document that is already open in the Executor's session. Never
104
+ edit an order in flight — issue a new annex, and say at the top which item it replaces.
105
+
106
+ `auditkit init <dir>` scaffolds the first three as empty templates plus an empty
107
+ `annexes/` directory for the fourth.
108
+
109
+ ### The deferred items ledger (in the plan of record)
110
+
111
+ A single running table, not prose scattered across the log:
112
+
113
+ ```
114
+ | ID | What | Deferred to | Closed by |
115
+ |---|---|---|---|
116
+ | D14 | Move the export job into the reporting module | P5 | — |
117
+ ```
118
+
119
+ Every decision (or annex) that pushes work to a future phase gets a row here the moment
120
+ it's made, in the same edit. **Before drafting any phase's tasks, read this table for
121
+ rows whose "Deferred to" matches the phase being drafted** — not a grep of the whole log
122
+ from memory. A row stays open until a task ID appears in "Closed by." An open row for a
123
+ phase that's about to be marked done is a stop, not a note for later: a promise made in
124
+ an early phase and never carried into the later phase's own task expansion is easy to
125
+ lose track of across a long run, and expensive to recover once several phases have
126
+ already closed on top of it. The ledger exists so that promise is a row someone has to
127
+ close, not a sentence someone has to remember.
128
+
129
+ ### The reserved-to-Auditor steps list (in the execution guide)
130
+
131
+ A running list, next to the rules of engagement, not a per-task Observations note:
132
+
133
+ ```
134
+ Steps reserved to the Auditor (never the Executor's model tier, whatever it is):
135
+ - The credential-adjacent check after a ship task (blocked for every model tier so far).
136
+ - Any commit touching a shared or generated resource (harness blocks it for every tier).
137
+ ```
138
+
139
+ The first time a harness or permission block turns out predictable — same step, blocked
140
+ for every Executor tried — add it here once. Every subsequent handoff whose task includes
141
+ that step cites this list instead of the step being rediscovered as a fresh `BLOCKED`
142
+ each time. A note written into one task's Observations after the first occurrence is read
143
+ once and never again; this list is read on every handoff.
144
+
145
+ ## Rules of engagement (the Executor follows these)
146
+
147
+ Put these at the head of the execution guide, ordered by how often they get broken.
148
+ Items 1-7 are fixed — copy them as they are. **Item 8 is not written from memory or
149
+ inherited from a different run; it is the first concrete step of starting a new run,
150
+ done before task 1 is drafted:**
151
+
152
+ - Check whether this codebase or team already states these somewhere authoritative — a
153
+ CONTRIBUTING file, CI config, an existing style guide, a prior run's execution guide.
154
+ If so, cite it; do not restate it.
155
+ - If nothing authoritative exists, **look at the project before asking** — grep for
156
+ what's already there rather than handing the user a blank checklist:
157
+
158
+ | Rule | Look for |
159
+ |---|---|
160
+ | Version bump per change | A version field (`package.json`, `pyproject.toml`, a `VERSION` file) and any script that already bumps it |
161
+ | No AI co-author trailers | `git log --grep="Co-Authored-By"` — existing trailers (or their consistent absence) |
162
+ | Fetch-before-push | Multiple contributors or branches in `git log`/`git branch -a` — a solo, single-branch repo needs this less |
163
+ | Quality gate green before push | CI config (`.github/workflows/`, `.gitlab-ci.yml`) or a `test`/`lint`/`build` script already wired together |
164
+ | i18n/locale parity | A locales/translations directory with more than one language file |
165
+ | Dependency lockstep | A lockfile and how strict its existing commit history is about touching it |
166
+
167
+ **Then ask with the finding attached, not a blank menu** — turn each hit into a
168
+ yes/no the user can answer in one word instead of a checklist they have to reason
169
+ through from nothing: *"Found `package.json` version `0.2.1` and no bump script — want
170
+ every task to bump it, and if so, how (semver rule, or you decide per task)?"* /
171
+ *"No CI config found — is there a command that should gate every push, or does this
172
+ repo not have one yet?"* This matters most for a user who isn't deep in the codebase
173
+ themselves (vibe-coding a project, not maintaining one they know by heart) — inspecting
174
+ first is the difference between a real choice and homework.
175
+ A rule with nothing detected for it is still offered, just without a finding attached.
176
+ - **Custom rules stay first-class** — a rule is usable as item 8 once it names what it
177
+ enforces, how it's checked, and what triggers it, whether or not it came from the
178
+ table above. Example: *"No schema migration touches a table over 10k rows without a
179
+ stated rollback plan in the task text, and the migration is run once with `--dry-run`
180
+ before it lands for real."* Names the trigger (a migration over the row threshold),
181
+ the requirement (a rollback plan, in the task text), and the check (`--dry-run` first).
182
+ - Settle it once, at the start of the run. Item 8 does not change task to task.
183
+
184
+ 1. **One task at a time, in order.** No batching. No starting a phase whose predecessor
185
+ is not `APPROVED`.
186
+ 2. **Never invent a value the codebase, infra, or environment declares.** Ports, origins,
187
+ domains, permission keys, versions, enum members — read them. Needing to ask for one
188
+ means a file was skipped.
189
+ 3. **Never guess a cause. Observe it.** Read the log, run the query, print the value.
190
+ When investigating a claim about the code, prefer whatever fast navigation tool this
191
+ project already has over blind grep — a code graph, ctags, an LSP, anything that
192
+ answers "who else reads this" faster than a full-text search. Never assume one is
193
+ installed or working; grep is the guaranteed fallback, not a default to reach past a
194
+ faster tool for. This check is quiet, not a topic — look, then use whatever's there.
195
+ Finding nothing is not itself worth telling the user; only bring it up if a tool was
196
+ found and is why the next command looks unusual.
197
+ 4. **Verification is running the thing, not reading the code.** "I reviewed it and it
198
+ looks correct" is reported as `FAILED`.
199
+ 5. **Scope is the task text.** Note unrelated problems in *Observations*; do not fix
200
+ them.
201
+ 6. **Stop and ask** when the instruction contradicts the code, when a task needs a
202
+ decision the document does not make, or when a previously passing check starts
203
+ failing for reasons unrelated to the change. An empty "Blocked" section across a
204
+ whole run reads as ambiguities resolved silently, not as ambiguities that never
205
+ existed. If you resolve one without stopping, log the decision.
206
+ 7. **When the system cannot do what a task asks, that is the deliverable.** Report it
207
+ and stop. A task that ends in a well-argued finding is a success.
208
+ 8. Project-specific rules go here: version bumping, commit attribution, dependency
209
+ sync before pushing, i18n parity, design-system checks, whatever this codebase
210
+ already enforces elsewhere. Cite the source file instead of restating it — one
211
+ source of truth. **Settled above, before task 1 — not filled in retroactively.**
212
+
213
+ ## Before expanding a phase into tasks
214
+
215
+ Read the deferred items ledger (above) for rows whose "Deferred to" matches the phase
216
+ about to be drafted. Every open row becomes a task, or gets re-deferred by editing the
217
+ row — never dropped silently. This replaces re-deriving the phase's obligations from
218
+ memory or a fresh grep of the whole log; the ledger is what makes that unnecessary.
219
+
220
+ ## Reporting back (Executor)
221
+
222
+ A compliance log entry is not addressed to anyone — it is a record. When a task, or the
223
+ last task of a phase, is done, close with a message back to the Auditor, same reasoning
224
+ as the handoff that started the work: whoever picks up the audit may be a fresh session
225
+ too, with nothing but the log to go on unless this exists.
226
+
227
+ ```
228
+ ROLE: Executor, reporting on <TASK-ID or "Phase <n>, tasks <first>-<last>">.
229
+
230
+ STATUS: <DONE | BLOCKED | FAILED> — logged in <COMPLIANCE_LOG>, section(s) already
231
+ filled in with literal command output.
232
+
233
+ Verify results: <the one-line summary an Auditor would want before deciding whether to
234
+ re-run everything themselves — not a substitute for that re-run>.
235
+
236
+ Deviations from the task text: <any, with the reasoning — or "none">.
237
+
238
+ Stops logged: <any "Blocked" entries raised mid-task per rule 6 — or "none">.
239
+
240
+ Ready for: <"Audit of Phase <n>" | "the next task, <ID>, once this is reviewed">.
241
+ ```
242
+
243
+ Do not narrate confidence ("this should be solid now") in place of the verify results —
244
+ the Auditor is about to re-run everything regardless (rule 4 applies to them too); a
245
+ report's job is to point at the evidence, not to argue for a verdict.
246
+
247
+ ## Auditing (Auditor)
248
+
249
+ **Re-run. Do not read the report and agree with it.**
250
+
251
+ 1. Run every suite yourself, with the project's own commands. A wrong invocation
252
+ produces a false failure and destroys your credibility for the rest of the audit —
253
+ confirm you're running from the right directory, against the right environment,
254
+ before treating an error as a finding.
255
+ 2. Re-execute the negative controls from scratch. Back the file up before mutating it
256
+ and restore from that copy — never with a source-control command that could discard
257
+ other uncommitted work in the same tree.
258
+ 3. Check artifacts exist: migrations, screenshots, specs, generated bundles.
259
+ 4. Read the checks, not just their names. Ask what ordering or fixture would make a
260
+ passing check pass for the wrong reason.
261
+ 5. **Write a probe when a claim is load-bearing.** A temporary check that asserts the
262
+ opposite of what the delivered one asserts tells you in thirty seconds whether the
263
+ behavior is real. Delete it afterward.
264
+ 6. Check timestamps before concluding something is missing. Work may have landed after
265
+ you looked.
266
+ 7. **A delivered negative control only proves the failure mode it was built to catch.**
267
+ Before signing a task off, ask what other claim in the same delivery is load-bearing
268
+ and has no probe of its own — do not accept the one control that exists as coverage
269
+ for the whole feature.
270
+ 8. **Before sending the verdict, check what it's missing, not just what it says.** This
271
+ applies to every verdict, not only a full-phase `APPROVED` or a `CONDITIONAL`/
272
+ `REJECTED` with an obvious remediation — a plain `PASS` on one task, mid-phase, with
273
+ something else already outstanding (an earlier remediation, the next task) needs the
274
+ same discipline: it does not end the message on its own. If anything is left to do —
275
+ a remediation handoff, the next task's handoff, or a re-statement of a handoff already
276
+ sent but not yet acted on — it ships in the same message as the verdict, as the
277
+ literal pasteable block from `references/handoffs.md`, not a sentence describing that
278
+ it's still pending. A verdict is not the deliverable; see "Remediation handoff" in
279
+ `references/handoffs.md`. In Autonomous mode, "ships in the same message" means the
280
+ Auditor launches that subagent in the same turn.
281
+
282
+ ### Verdicts
283
+
284
+ | Verdict | Meaning |
285
+ |---|---|
286
+ | `APPROVED` | Every gate independently verified |
287
+ | `CONDITIONAL` | Accepted except for named items; state which and what closes them |
288
+ | `REJECTED` | The deliverable does not do what it claims, even if it is green |
289
+
290
+ A passing check that documents wrong behavior as correct is a **rejection**, not an
291
+ observation. It will defend the defect against whoever tries to fix it later.
292
+
293
+ Separate **process findings** from **code findings**. Work can be accepted on its
294
+ merits while the control that should have caught a defect is recorded as failed. Good
295
+ outcomes do not validate a broken process backward. A condition holds as written or it
296
+ stays open — do not fold an unmet condition into "recorded, not blocking" just to avoid
297
+ holding up a phase; that turns a defect into paperwork. Process findings include: the
298
+ wrong model tier (or an inherited one), a small-model task that improvised instead of
299
+ escalating, a negative control restored with a command that discarded other uncommitted
300
+ work, and a question sent to the owner that belonged to the Auditor.
301
+
302
+ ### Closing the run
303
+
304
+ A verdict says whether one phase or task passed. It does not say whether the *run* is
305
+ over — those are different claims, and leaving the second one to tone or inference
306
+ means the user ends up asking "so is that everything?" directly, which is the signal
307
+ this section exists to make unnecessary.
308
+
309
+ Before saying a run is over, check the plan of record itself for a phase that exists on
310
+ paper but was never turned into tasks — not just for a next task in the current
311
+ execution guide. A phase can be planned and never scheduled; "nothing left in the guide
312
+ I'm looking at" and "nothing left in the plan" are not the same check.
313
+
314
+ When every phase is `APPROVED` and the plan of record declares no further phase, say so
315
+ as its own line, not folded into other prose:
316
+
317
+ ```
318
+ RUN COMPLETE — <plan of record name>. Every phase APPROVED. No further phase declared.
319
+ Nothing outstanding.
320
+ ```
321
+
322
+ Anything short of that — one phase closing, a conditional approval, work still queued —
323
+ uses the verdict-plus-handoff shapes in `references/handoffs.md` instead. `RUN COMPLETE` is reserved for the
324
+ one message that actually ends the need for another Auditor turn on this plan of record.
325
+
326
+ ## The Auditor is bound by rule 4 too
327
+
328
+ An order can be wrong in the same way a delivery can be wrong, and it is more
329
+ dangerous when it is: the Executor is told a decided item is a stop, not a choice, so a
330
+ wrong decided item arrives armored against the one person positioned to catch it.
331
+
332
+ The failure has one shape: the Auditor asserted a fact about the system — a field
333
+ exists, a function behaves a certain way, a value has a certain sign — from reasoning
334
+ or a single grep hit, without running the command that would actually confirm it.
335
+ **Deliveries get audited by running them; orders too often get written by reasoning
336
+ about them.** That asymmetry is the whole problem.
337
+
338
+ Two rules close it:
339
+
340
+ **1. An order may not assert what it has not run.** Any claim in an order about a
341
+ field, a function's behavior, or a runtime value carries the command that established
342
+ it, pasted, in the order. Rule 4 above — *verification is running the thing, not
343
+ reading the code* — is not only the Executor's rule. It is the protocol's. If the
344
+ Auditor cannot paste the command, the Auditor cannot assert the fact; it goes in the
345
+ order as a question for the Executor to establish, not as a decision.
346
+
347
+ For anything touching money, permissions, or schema, the cheapest form of this is a
348
+ falsifying probe: write the smallest script that would prove the claim wrong and run
349
+ it before the order goes out.
350
+
351
+ **2. A decided item is not reopenable; a factual claim always is.**
352
+
353
+ | | Reopenable by the Executor? |
354
+ |---|---|
355
+ | A decision — scope, trade-off, design | **No.** Believing otherwise is a stop. |
356
+ | A factual claim — this field exists, this function does X, this value is Y | **Always**, with a pasted command that contradicts it. The evidence outranks the order. |
357
+
358
+ An Executor who runs a command that falsifies an order is not exceeding scope; that is
359
+ the second pair of eyes the arrangement exists to buy.
360
+
361
+ ## Maintaining the log
362
+
363
+ A compliance log that only ever gets appended to accumulates duplicate paragraphs —
364
+ the same open condition re-stated verbatim across several entries because it was
365
+ easier to copy the last summary than to write a new one. That is a cost, not a
366
+ neutral habit: it makes the current state of the run more expensive to find for
367
+ whoever reads it next, human or agent. Compact the log periodically — collapse a
368
+ condition that has appeared unchanged across several entries into one current
369
+ statement with a pointer to when it opened. `auditkit lint` flags near-duplicate
370
+ paragraphs so this doesn't have to be caught by eye.
371
+
372
+ **Annex count is a health signal, not just a history.** A phase that accumulates many
373
+ corrections after work started is a phase whose plan was under-verified before work
374
+ started, more often than it is a phase that hit genuine surprises. If one phase is
375
+ generating annexes faster than the others, that is worth naming as a finding about the
376
+ planning step, not only fixing task by task. `auditkit lint` warns past a configurable
377
+ threshold.
378
+
379
+ ## Repo conventions this rides on
380
+
381
+ Cite whatever this codebase already enforces — version bumping, commit attribution,
382
+ fetch-before-push, environment-value discipline, evidence-driven debugging — from its
383
+ own source of truth rather than restating it here, so there's one place it can drift
384
+ out of sync from.
385
+
386
+ ## The CLI
387
+
388
+ If `auditkit` is installed, prefer it over doing these by hand:
389
+
390
+ - `auditkit init <dir>` — scaffold the plan of record, execution guide, and compliance
391
+ log from templates, plus an empty `annexes/` directory.
392
+ - `auditkit lint <dir>` — cross-check task IDs between the execution guide and the
393
+ compliance log in both directions, flag `DONE` reports with no pasted verify output
394
+ (a lone "n/a" or "TBD" is not output),
395
+ flag near-duplicate paragraphs, warn on annex count per phase, flag gates with no
396
+ stated negative control (it must name the action that removes the protection and the
397
+ failure that follows; a negated, deferred, or waived mention such as "n/a" or "TBD" does
398
+ not count). Missing documents, an execution guide with no tasks, and tasks without a
399
+ `**Report:**` line are errors, not a clean result. Lint checks that evidence is
400
+ present, not that it is authentic: re-running it is still the Auditor's job.
401
+ - `auditkit negcontrol --file <path> --break-cmd "<cmd>" --test-cmd "<cmd>"` — backs the
402
+ file up, runs the break command, runs the test (expects failure), restores from the
403
+ backup, runs the test again (expects success), and prints a paste-ready transcript.
404
+ With `--restore-cmd`, the file must still end byte-identical to the backup or it is
405
+ restored from it; `--timeout` bounds each command.
406
+ - `auditkit status <dir>` — combines the status board verdicts with the latest report
407
+ per ID and prints what's still open. Only `APPROVED` closes an item; a `DONE` with no
408
+ verdict is listed as awaiting audit.
409
+
410
+ See the repo's `README.md` for install instructions.