agentic-sdd-framework 1.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/AGENTS.template.md +59 -0
- package/.agents/CONTEXT.template.md +41 -0
- package/.agents/ENTRYPOINT.template.md +31 -0
- package/.agents/skills/ast-navigator/SKILL.md +31 -0
- package/.agents/skills/ast-navigator/adapters/ast-grep.md +18 -0
- package/.agents/skills/ast-navigator/adapters/graphify.md +19 -0
- package/.agents/skills/ast-navigator/adapters/lsp.md +16 -0
- package/.agents/skills/ast-navigator/adapters/ripgrep.md +19 -0
- package/.agents/skills/auditor-executor-protocol/SKILL.md +410 -0
- package/.agents/skills/auditor-executor-protocol/references/autonomous-mode.md +144 -0
- package/.agents/skills/auditor-executor-protocol/references/failure-modes-and-example.md +103 -0
- package/.agents/skills/auditor-executor-protocol/references/handoffs.md +133 -0
- package/.agents/skills/auditor-executor-protocol/references/tasks-and-gates.md +81 -0
- package/.agents/skills/no-ai-slop/LICENSE +21 -0
- package/.agents/skills/no-ai-slop/SKILL.md +52 -0
- package/.agents/skills/strategic-cto/SKILL.md +54 -0
- package/CHANGELOG.md +117 -0
- package/LICENSE +21 -0
- package/README.md +244 -0
- package/docs/SPEC_TEMPLATE.md +78 -0
- package/docs/decisions/ADR_TEMPLATE.md +49 -0
- package/docs/guidelines/AST_NAVIGATION.md +51 -0
- package/docs/guides/AGENT_CREDENTIALS.md +75 -0
- package/docs/guides/GITHUB_CLI_SETUP.md +74 -0
- package/docs/incidents/0000-00-00-incident-template.md +35 -0
- package/docs/roadmap/templates/compliance-log.md +37 -0
- package/docs/roadmap/templates/execution-guide.md +75 -0
- package/docs/roadmap/templates/plan-of-record.md +49 -0
- package/package.json +49 -0
- package/scripts/check-copy-slop.js +120 -0
- package/scripts/check-file-size.js +66 -0
- package/scripts/check-spec.js +201 -0
- package/scripts/check-system-prerequisites.js +133 -0
- package/scripts/check-versions.js +50 -0
- package/scripts/dev/fuzz-spec-markup.js +123 -0
- package/scripts/dev/set-npm-publish-token.sh +40 -0
- package/scripts/dev/sync-vendored.js +94 -0
- package/scripts/install-git-hooks.js +103 -0
- package/scripts/lib/cli.js +60 -0
- package/scripts/lib/config.js +111 -0
- package/scripts/lib/git.js +211 -0
- package/scripts/lib/markdown.js +46 -0
- package/scripts/lib/provision.js +323 -0
- package/scripts/lib/runner.js +70 -0
- package/scripts/lib/sdd.config.schema.json +213 -0
- package/scripts/lib/slop-patterns.js +57 -0
- package/scripts/lib/spec-markup.js +346 -0
- package/scripts/lib/spec.js +226 -0
- package/scripts/lib/state.js +107 -0
- package/scripts/lib/vendor/README.md +11 -0
- package/scripts/lib/vendor/markdown-it.LICENSE +22 -0
- package/scripts/lib/vendor/markdown-it.min.js +3 -0
- package/scripts/quality-gate.js +151 -0
- package/scripts/sdd-init.js +245 -0
- package/scripts/sdd-verify.js +176 -0
- package/scripts/verify-no-secrets.js +216 -0
- package/sdd.config.json +33 -0
|
@@ -0,0 +1,59 @@
|
|
|
1
|
+
# Agent Constitution and Behavioral Guidelines
|
|
2
|
+
|
|
3
|
+
This document establishes the non-negotiable operating rules for AI coding agents working in this repository. Every agent must read and adhere to these directives before executing tasks.
|
|
4
|
+
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
## 1. Discovery First (No Premature Assumptions)
|
|
8
|
+
* **Rule:** Before recommending architectures, selecting frameworks, or generating code on a new initiative, the agent must execute the 4-Pillar Discovery Interview (Scale/Concurrency, Hardware/Deployment, Workload/Compute, Modularity).
|
|
9
|
+
* **Why this rule exists:**
|
|
10
|
+
> [Document the incident or rationale here. For example: An agent previously recommended a heavy full-stack framework for a lightweight streaming utility, introducing unnecessary runtime bloat.]
|
|
11
|
+
|
|
12
|
+
---
|
|
13
|
+
|
|
14
|
+
## 2. Evidence-Driven Debugging & Diagnostics
|
|
15
|
+
* **Rule:** Never guess root causes or apply speculative patches. Inspect log files, inspect command output, and run diagnostics before altering code.
|
|
16
|
+
* **Why this rule exists:**
|
|
17
|
+
> [Document the incident or rationale here. For example: Speculative patches often introduce regression cascades that obscure the original defect.]
|
|
18
|
+
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
## 3. Mandatory Verification Before Certification
|
|
22
|
+
* **Rule:** A task or phase is not complete until its explicit verification command exits with code 0. Reading code visually is never a substitute for running the code.
|
|
23
|
+
* **Why this rule exists:**
|
|
24
|
+
> [Document the incident or rationale here. For example: Clean syntax does not guarantee functional runtime behavior or integration success.]
|
|
25
|
+
|
|
26
|
+
---
|
|
27
|
+
|
|
28
|
+
## 4. Closed-Network Testing Isolation
|
|
29
|
+
* **Rule:** Automated test suites must never contact external internet hosts. All external integrations must be mocked or gated on environment variables. Loopback testing is permitted for local servers.
|
|
30
|
+
* **Why this rule exists:**
|
|
31
|
+
> [Document the incident or rationale here. For example: Test suites that hit third-party APIs can cause hard bounces, rate limits, and unexpected billing exposure.]
|
|
32
|
+
|
|
33
|
+
---
|
|
34
|
+
|
|
35
|
+
## 5. Zero-Trust Secrets Management
|
|
36
|
+
* **Rule:** Agents must never request API keys or credentials in chat prompts. Secrets must be read directly from the Tier 3 Vault (`~/secrets/<app>/.vault`) or environment variables. Never commit secrets to Git. Tier model: `docs/guides/AGENT_CREDENTIALS.md`.
|
|
37
|
+
* **Why this rule exists:**
|
|
38
|
+
> [Document the incident or rationale here. For example: Chat logs and git histories are permanently recorded; leaking credentials requires immediate key revocation.]
|
|
39
|
+
|
|
40
|
+
---
|
|
41
|
+
|
|
42
|
+
## 6. Scope Bounding & Atomic Progression
|
|
43
|
+
* **Rule:** Execute one task at a time in strict sequence. Do not refactor unrelated files or perform out-of-scope cleanups without explicit Auditor authorization.
|
|
44
|
+
* **Why this rule exists:**
|
|
45
|
+
> [Document the incident or rationale here. For example: Unbounded edits make regression tracing difficult and invalidate concurrent session handoffs.]
|
|
46
|
+
|
|
47
|
+
---
|
|
48
|
+
|
|
49
|
+
## 7. Factual Technical Copy (No AI Slop)
|
|
50
|
+
* **Rule:** All documentation, user-facing copy, and commit messages must be factual, concise, and dense, following the `no-ai-slop` skill. The quality gate enforces its banned patterns.
|
|
51
|
+
* **Why this rule exists:**
|
|
52
|
+
> [Document the incident or rationale here. For example: Promotional puffery and marketing contrasts obscure technical reality and degrade documentation clarity.]
|
|
53
|
+
|
|
54
|
+
---
|
|
55
|
+
|
|
56
|
+
## 8. Author Attribution & Integrity
|
|
57
|
+
* **Rule:** Never attach `Co-Authored-By:` trailers crediting AI assistants to Git commits. The repository owner is the sole author.
|
|
58
|
+
* **Why this rule exists:**
|
|
59
|
+
> [Document the incident or rationale here. For example: Project contributors retain 100% ownership and commit graph clarity without AI tooling noise.]
|
|
@@ -0,0 +1,41 @@
|
|
|
1
|
+
# Operational Context and Memory Registry
|
|
2
|
+
|
|
3
|
+
This file records the cumulative architectural decisions, production incidents, and operational learnings for this repository. AI agents read this file to avoid repeating past mistakes.
|
|
4
|
+
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
## 1. Project Summary & Active Capabilities
|
|
8
|
+
|
|
9
|
+
* **Project Name:** [Project Name]
|
|
10
|
+
* **Specification Mode:** Lite | Rigor
|
|
11
|
+
* **Primary Runtime:** [e.g. Node.js 24 LTS, Go 1.23, Python 3.12]
|
|
12
|
+
* **Architecture Style:** Clean Architecture | Standard MVC | Modular Monolith
|
|
13
|
+
* **Active Capabilities:**
|
|
14
|
+
* No-AI-Slop Linter: [Enabled / Disabled]
|
|
15
|
+
* AST Navigator Adapter: [graphify / ast-grep / ripgrep / lsp]
|
|
16
|
+
* Internationalization (i18n): [Enabled / Disabled]
|
|
17
|
+
* Progressive Web App (PWA): [Enabled / Disabled]
|
|
18
|
+
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
## 2. Incident & Root-Cause Registry
|
|
22
|
+
|
|
23
|
+
Whenever a production regression, test flake, or operational issue is resolved, record the incident below. Every incident must produce a verifiable rule in `.agents/AGENTS.md`.
|
|
24
|
+
|
|
25
|
+
### Template:
|
|
26
|
+
```markdown
|
|
27
|
+
### [YYYY-MM-DD] Incident: [Short Title]
|
|
28
|
+
* **Symptom:** What broke and where was it observed.
|
|
29
|
+
* **Root Cause:** The underlying mechanical or architectural flaw.
|
|
30
|
+
* **Remediation:** The code change or configuration fix applied.
|
|
31
|
+
* **Inviolable Rule Created:** Link to the corresponding rule added to AGENTS.md.
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
---
|
|
35
|
+
|
|
36
|
+
## 3. Active Technical Debt & Non-Goals
|
|
37
|
+
|
|
38
|
+
Document known constraints, deferred refactors, and explicit non-goals to prevent agents from attempting unsolicited cleanups.
|
|
39
|
+
|
|
40
|
+
* [Non-Goal 1]: [Description]
|
|
41
|
+
* [Technical Debt 1]: [Description and tracking issue]
|
|
@@ -0,0 +1,31 @@
|
|
|
1
|
+
<!-- sdd:managed. Generated by sdd-init; rerun it to refresh. Put project rules in .agents/AGENTS.md. -->
|
|
2
|
+
# Agent Entry Point: {{PROJECT_NAME}}
|
|
3
|
+
|
|
4
|
+
This project follows the Agentic SDD Framework. Before any task, read:
|
|
5
|
+
|
|
6
|
+
1. `.agents/AGENTS.md`: non-negotiable rules for this repository.
|
|
7
|
+
2. `.agents/CONTEXT.md`: incident registry, technical debt, and non-goals.
|
|
8
|
+
3. {{SPEC_LOCATION}}
|
|
9
|
+
|
|
10
|
+
## Skills
|
|
11
|
+
|
|
12
|
+
Load a skill from `.agents/skills/<name>/SKILL.md` only when its trigger applies.
|
|
13
|
+
|
|
14
|
+
| Skill | Load when |
|
|
15
|
+
| :--- | :--- |
|
|
16
|
+
| `strategic-cto` | Choosing a stack or architecture, or answering "can this be improved?" |
|
|
17
|
+
| `auditor-executor-protocol` | Work spans several phases or sessions, or another agent implements the plan |
|
|
18
|
+
| `no-ai-slop` | Writing documentation, user-facing copy, or commit messages |
|
|
19
|
+
| `ast-navigator` | Exploring code. Active adapter: `{{AST_ADAPTER}}` (`.agents/skills/ast-navigator/adapters/{{AST_ADAPTER}}.md`) |
|
|
20
|
+
|
|
21
|
+
## Project Facts
|
|
22
|
+
|
|
23
|
+
* Runtime: {{RUNTIME}}
|
|
24
|
+
* Specification mode: {{SPEC_MODE}}
|
|
25
|
+
* Configuration: `sdd.config.json`
|
|
26
|
+
|
|
27
|
+
## Before Every Push
|
|
28
|
+
|
|
29
|
+
Run `{{GATE_COMMAND}}`. It must exit 0. The pre-push hook runs it automatically, and it checks the specification as well as the code: a checked task or `DONE` report without pasted evidence fails.
|
|
30
|
+
|
|
31
|
+
{{SPEC_CHECK}}
|
|
@@ -0,0 +1,31 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: ast-navigator
|
|
3
|
+
description: "MANDATORY for codebase exploration, symbol lookups, and dependency tracing. Enforces AST-first structural navigation before falling back to full-text grep, minimizing token usage and context clutter. Supports pluggable adapters: Graphify, ast-grep, ripgrep, and LSP/SCIP."
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# AST Navigator Protocol: Token-Efficient Code Exploration
|
|
7
|
+
|
|
8
|
+
## Objective
|
|
9
|
+
Enable AI coding agents to navigate complex codebases by querying Abstract Syntax Tree (AST) representations and dependency graphs first, drastically reducing LLM context consumption compared to reading entire files or running unbounded text searches.
|
|
10
|
+
|
|
11
|
+
---
|
|
12
|
+
|
|
13
|
+
## 🧭 Navigation Hierarchy
|
|
14
|
+
|
|
15
|
+
1. **Tier 1 (AST / Graph Query):** Query structural dependencies, class hierarchies, and symbol definitions using the configured adapter in `sdd.config.json`.
|
|
16
|
+
2. **Tier 2 (Fallback to Grep):** Fall back to targeted regex grep only when:
|
|
17
|
+
* The AST adapter returns no matches.
|
|
18
|
+
* The graph index is known to be stale and cannot be updated immediately.
|
|
19
|
+
* Exact line-level character sequences (string literals, CSS rules, comments) are required.
|
|
20
|
+
3. **Tier 3 (File Inspection):** Read specific line slices (`view_file`) only after locating the exact symbol and line numbers. Never dump entire 1,000-line files into context without target line bounds.
|
|
21
|
+
|
|
22
|
+
---
|
|
23
|
+
|
|
24
|
+
## 🔌 Active Adapter
|
|
25
|
+
|
|
26
|
+
Read `capabilities.astNavigation.adapter` in `sdd.config.json`, then load only `adapters/<adapter>.md` for its commands. Available adapters: `ast-grep`, `graphify`, `ripgrep`, `lsp`. The comparison used to choose one lives in `docs/guidelines/AST_NAVIGATION.md`.
|
|
27
|
+
|
|
28
|
+
---
|
|
29
|
+
|
|
30
|
+
## 🚨 Verification Rule Before Claiming Compliance
|
|
31
|
+
A missing edge or symbol in an AST index is **not** definitive proof that a dependency does not exist. Before certifying architectural compliance or absence of a vulnerability, always verify the actual import statements of the target file.
|
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
# Adapter: ast-grep
|
|
2
|
+
|
|
3
|
+
* **Upstream Creator:** Herrington Darkholme (`@HerringtonDarkholme`)
|
|
4
|
+
* **Repository:** [https://github.com/ast-grep/ast-grep](https://github.com/ast-grep/ast-grep)
|
|
5
|
+
* **License:** MIT License
|
|
6
|
+
* **Runtime Requirement:** Native binary (Rust-powered, installable via `npm i -g @ast-grep/cli` or `brew install ast-grep`)
|
|
7
|
+
|
|
8
|
+
## Capabilities and Mechanics
|
|
9
|
+
ast-grep performs structural code searches and transformations using Tree-sitter syntax trees. Unlike plain regex, it matches syntax constructs regardless of formatting, indentation, or comments.
|
|
10
|
+
|
|
11
|
+
## Common Operational Commands
|
|
12
|
+
* **Search AST Pattern:** `sg run -p '$FUNCTION($$$ARGS)' -l typescript`
|
|
13
|
+
* **Find Class Definitions:** `sg run -p 'class $NAME extends $BASE { $$$ }'`
|
|
14
|
+
* **Scan Rules:** `sg scan` (executes pre-configured YAML linting rules).
|
|
15
|
+
|
|
16
|
+
## Best Practices
|
|
17
|
+
* Use for lightweight environments where Python dependencies are undesirable.
|
|
18
|
+
* Ideal for CI/CD structural rule enforcement and automated AST refactors.
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
# Adapter: Graphify
|
|
2
|
+
|
|
3
|
+
* **Upstream Organization:** Graphify Labs (`@Graphify-Labs`)
|
|
4
|
+
* **Repository:** [https://github.com/Graphify-Labs/graphify](https://github.com/Graphify-Labs/graphify)
|
|
5
|
+
* **License:** Apache License 2.0
|
|
6
|
+
* **Runtime Requirement:** Python 3.10+ (`pip install graphifyy`)
|
|
7
|
+
|
|
8
|
+
## Capabilities and Mechanics
|
|
9
|
+
Graphify extracts relational knowledge graphs of symbols, imports, and calls across multi-language repositories, generating interactive visual graphs and structured JSON query outputs.
|
|
10
|
+
|
|
11
|
+
## Common Operational Commands
|
|
12
|
+
* **Update Graph:** `graphify update .` (or configured project script `npm run graph:update`).
|
|
13
|
+
* **Check Index Freshness:** `graphify check` (validates graph commit against git HEAD).
|
|
14
|
+
* **Query Symbol Path:** `graphify path <source-node> <target-node>`.
|
|
15
|
+
* **Explain Subgraph:** `graphify explain <symbol-name>`.
|
|
16
|
+
|
|
17
|
+
## Best Practices
|
|
18
|
+
* Always regenerate the graph after major refactors or file deletions.
|
|
19
|
+
* Exclude build noise directories (`dist/`, `build/`, `.venv/`) from the index.
|
|
@@ -0,0 +1,16 @@
|
|
|
1
|
+
# Adapter: Language Server Protocol (LSP) / SCIP
|
|
2
|
+
|
|
3
|
+
* **Upstream Organization:** SCIP Code (`@scip-code`), originally created by Sourcegraph
|
|
4
|
+
* **Repository:** [https://github.com/scip-code/scip](https://github.com/scip-code/scip)
|
|
5
|
+
* **License:** Apache License 2.0
|
|
6
|
+
* **Runtime Requirement:** Language-specific indexers (`scip-typescript`, `scip-go`, etc.)
|
|
7
|
+
|
|
8
|
+
## Capabilities and Mechanics
|
|
9
|
+
SCIP (Source Code Intelligence Protocol) indexes definitions, references, and hover types across strongly typed codebases. It enables exact compiler-level cross-file navigation without heuristics.
|
|
10
|
+
|
|
11
|
+
## Common Operational Commands
|
|
12
|
+
* **Index Project:** `scip-typescript index`
|
|
13
|
+
* **Query References:** Query through SCIP CLI or LSP client interfaces.
|
|
14
|
+
|
|
15
|
+
## Best Practices
|
|
16
|
+
* Use in large enterprise repositories where type relationships and cross-package references must be verified with compiler precision.
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
# Adapter: ripgrep
|
|
2
|
+
|
|
3
|
+
* **Upstream Creator:** Andrew Gallant (`@BurntSushi`)
|
|
4
|
+
* **Repository:** [https://github.com/BurntSushi/ripgrep](https://github.com/BurntSushi/ripgrep)
|
|
5
|
+
* **License:** MIT License / UNLICENSE
|
|
6
|
+
* **Runtime Requirement:** Native binary (`rg`)
|
|
7
|
+
|
|
8
|
+
## Capabilities and Mechanics
|
|
9
|
+
ripgrep is an ultra-fast line-oriented search tool that respects `.gitignore` rules by default. It provides the universal baseline fallback when AST parsers cannot resolve a query or when searching non-code files (markdown, JSON, YAML).
|
|
10
|
+
|
|
11
|
+
## Common Operational Commands
|
|
12
|
+
* **Find Literal String:** `rg -F "exact_symbol_name"`
|
|
13
|
+
* **Case-Insensitive Search:** `rg -i "pattern"`
|
|
14
|
+
* **Filter by File Extension:** `rg "pattern" -g "*.ts"`
|
|
15
|
+
* **List Matching Files Only:** `rg -l "pattern"`
|
|
16
|
+
|
|
17
|
+
## Best Practices
|
|
18
|
+
* Always use `--type` or `-g` globs to restrict search scope when querying large repositories.
|
|
19
|
+
* Avoid dumping thousands of lines into agent context; pair with `head` or line count limits.
|
|
@@ -0,0 +1,410 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: auditor-executor-protocol
|
|
3
|
+
description: "Runs multi-phase work as a two-role protocol: an Auditor who writes numbered, verifiable tasks and signs off phases, and an Executor who implements one task at a time and reports evidence. Use whenever a piece of work is large enough to span several phases or sessions, whenever handing a plan to another agent (subagent, spawned session, or a different tool) to implement, or when asked to audit/verify work another agent reports as complete. Also use when a plan needs to become executable instructions rather than a discussion document. Covers: the document set, rules of engagement, task and gate format, how to audit by re-running rather than reading, negative controls, verdict vocabulary, the annex pattern for correcting an order that is already in flight, and an optional operating mode chosen per run — Guided (the owner relays each handoff) or Autonomous (the Auditor launches Executors as subagents with an explicit per-task model tier, answers their BLOCKED questions itself, and interrupts the owner only for a numbered critical-risk list). Triggers: 'audit', 'auditor', 'executor', 'phased plan', 'exit criteria', 'gate', 'compliance log', 'remediation', 'handoff', 'have another agent implement this', 'verify what was reported', 'sign off', 'autonomous mode', 'guided mode'. NOT for reviewing a diff or a pull request for bugs — that's a code review; this protocol audits an execution report against pre-declared gates, and governs how multi-phase work is handed off and signed off."
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Auditor / Executor Protocol
|
|
7
|
+
|
|
8
|
+
Two roles, one paper trail. The Auditor decides what "done" means and proves it
|
|
9
|
+
independently. The Executor implements and reports evidence. Neither does the other's job.
|
|
10
|
+
|
|
11
|
+
This exists because a plan handed to an implementing agent is not instructions until
|
|
12
|
+
someone makes it unambiguous, and a report that says "all green" is not verification
|
|
13
|
+
until someone re-runs it. Both failures are common and both are expensive once they
|
|
14
|
+
compound across dozens of tasks.
|
|
15
|
+
|
|
16
|
+
A companion CLI, `auditkit`, operationalizes the mechanical parts of this: scaffolding
|
|
17
|
+
the document set, checking the task list and the log agree with each other, running a
|
|
18
|
+
negative control end to end, and printing where a run stands. Use it if it's installed;
|
|
19
|
+
follow this document either way.
|
|
20
|
+
|
|
21
|
+
## When to use this
|
|
22
|
+
|
|
23
|
+
- Work spanning more than one phase or more than one session.
|
|
24
|
+
- Handing a plan to another agent to implement — a subagent, a spawned session, a
|
|
25
|
+
different tool, another person.
|
|
26
|
+
- Auditing work someone else reports as complete.
|
|
27
|
+
- Turning a roadmap into something executable.
|
|
28
|
+
|
|
29
|
+
Not for single-session tasks you are doing yourself. The overhead only pays off when the
|
|
30
|
+
implementer is not the person who wrote the plan.
|
|
31
|
+
|
|
32
|
+
## The two roles
|
|
33
|
+
|
|
34
|
+
| | Auditor | Executor |
|
|
35
|
+
|---|---|---|
|
|
36
|
+
| Owns | The instructions, the gates, the verdicts | The implementation, the evidence |
|
|
37
|
+
| Never | Writes feature code | Redesigns, or decides scope |
|
|
38
|
+
| Output | Task expansions, verdicts, remediation orders | Working code, pasted command output |
|
|
39
|
+
|
|
40
|
+
Whoever writes the plan is the Auditor. Say which role you are holding at the start of a
|
|
41
|
+
session so it does not drift.
|
|
42
|
+
|
|
43
|
+
## Reference files (load on demand)
|
|
44
|
+
|
|
45
|
+
The core rules are in this file. Load a reference file only when you reach the step it
|
|
46
|
+
covers; each one is self-contained and refers back to sections of this file by name.
|
|
47
|
+
|
|
48
|
+
| File | Load when |
|
|
49
|
+
|---|---|
|
|
50
|
+
| [`references/tasks-and-gates.md`](references/tasks-and-gates.md) | Writing a task or a gate (Auditor) |
|
|
51
|
+
| [`references/handoffs.md`](references/handoffs.md) | Handing a task to an Executor, or writing a remediation after a verdict that is not a clean `APPROVED` (Auditor) |
|
|
52
|
+
| [`references/autonomous-mode.md`](references/autonomous-mode.md) | The run is in Autonomous mode, or you are writing the Auditor brief that starts one |
|
|
53
|
+
| [`references/failure-modes-and-example.md`](references/failure-modes-and-example.md) | Auditing a delivery, or you want to see a filled-in run end to end |
|
|
54
|
+
|
|
55
|
+
## Operating mode: Guided or Autonomous (chosen at the start of the run)
|
|
56
|
+
|
|
57
|
+
The mode is optional and the owner's choice. Two ways to run the protocol:
|
|
58
|
+
|
|
59
|
+
| | Guided mode | Autonomous mode |
|
|
60
|
+
|---|---|---|
|
|
61
|
+
| Who launches Executors | The owner pastes each handoff into a fresh session | The Auditor, as subagents |
|
|
62
|
+
| Who answers `BLOCKED` | The owner, or the Auditor through the owner | The Auditor, with a decision entry in the log |
|
|
63
|
+
| Owner's role | Relays every message, sees every verdict | Receives one short report per closed phase; interrupted only for the critical-risk list |
|
|
64
|
+
| Fits when | The owner wants to watch and approve each step, the work is new territory, or the Auditor has no subagent tool | The plan is settled and the owner wants it run end to end |
|
|
65
|
+
|
|
66
|
+
How the mode is chosen:
|
|
67
|
+
|
|
68
|
+
- **If the owner already named it** ("autonomous", "guided", or an unambiguous
|
|
69
|
+
equivalent like "don't interrupt me for every step"), use it. Do not ask again.
|
|
70
|
+
- **If not, ask once, before task 1**, as a single two-option question with a one-line
|
|
71
|
+
recommendation (Guided for a first run in unfamiliar code or anything touching money
|
|
72
|
+
or production data with no rehearsal; Autonomous for a settled plan with strong gates).
|
|
73
|
+
This is one of the few questions that is genuinely the owner's: it decides how often
|
|
74
|
+
they get interrupted for the rest of the run.
|
|
75
|
+
- **If the environment has no subagent tool**, Autonomous is not available; say so and
|
|
76
|
+
run Guided.
|
|
77
|
+
- Record the choice as a governance decision in the plan of record, next to the rules of
|
|
78
|
+
engagement. Switching modes mid-run is allowed only when the owner asks, and is
|
|
79
|
+
recorded by annex. The Auditor never switches itself from Guided to Autonomous; it may
|
|
80
|
+
drop from Autonomous to Guided only through an item on the critical-risk list.
|
|
81
|
+
|
|
82
|
+
Everything else in this document applies to both modes. Autonomous mode adds the rules in
|
|
83
|
+
`references/autonomous-mode.md`; it removes none. Re-running, negative controls and separate
|
|
84
|
+
process/code findings matter more when nobody is watching, not less.
|
|
85
|
+
|
|
86
|
+
## The document set
|
|
87
|
+
|
|
88
|
+
Four documents. Keep them separate; merging them is how the instructions turn back into
|
|
89
|
+
a discussion.
|
|
90
|
+
|
|
91
|
+
1. **Plan of record** — the *what* and the *why*. Phases, decisions, trade-offs. Nobody
|
|
92
|
+
implements from this. Carries the **deferred items ledger** (below).
|
|
93
|
+
2. **Execution guide** — the *how*. Numbered tasks (`P<phase>-T<n>`, or any scheme with a
|
|
94
|
+
stable, greppable ID), each with files, steps, a verification command, and its
|
|
95
|
+
expected output. Gates (`P<phase>-G<n>`) close each phase. Carries the
|
|
96
|
+
**reserved-to-Auditor steps** list (below).
|
|
97
|
+
3. **Compliance log** — where the Executor reports. Pre-generate one empty row per task
|
|
98
|
+
and gate ID so nothing can be quietly skipped.
|
|
99
|
+
4. **Remediation order (annex)** — written by the Auditor after an audit that isn't a
|
|
100
|
+
clean `APPROVED`. Self-contained: the Executor must not need the audit conversation
|
|
101
|
+
or the full log to act on it. Template and reasoning: "Remediation handoff" in `references/handoffs.md`.
|
|
102
|
+
|
|
103
|
+
**Annexes** supersede a document that is already open in the Executor's session. Never
|
|
104
|
+
edit an order in flight — issue a new annex, and say at the top which item it replaces.
|
|
105
|
+
|
|
106
|
+
`auditkit init <dir>` scaffolds the first three as empty templates plus an empty
|
|
107
|
+
`annexes/` directory for the fourth.
|
|
108
|
+
|
|
109
|
+
### The deferred items ledger (in the plan of record)
|
|
110
|
+
|
|
111
|
+
A single running table, not prose scattered across the log:
|
|
112
|
+
|
|
113
|
+
```
|
|
114
|
+
| ID | What | Deferred to | Closed by |
|
|
115
|
+
|---|---|---|---|
|
|
116
|
+
| D14 | Move the export job into the reporting module | P5 | — |
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
Every decision (or annex) that pushes work to a future phase gets a row here the moment
|
|
120
|
+
it's made, in the same edit. **Before drafting any phase's tasks, read this table for
|
|
121
|
+
rows whose "Deferred to" matches the phase being drafted** — not a grep of the whole log
|
|
122
|
+
from memory. A row stays open until a task ID appears in "Closed by." An open row for a
|
|
123
|
+
phase that's about to be marked done is a stop, not a note for later: a promise made in
|
|
124
|
+
an early phase and never carried into the later phase's own task expansion is easy to
|
|
125
|
+
lose track of across a long run, and expensive to recover once several phases have
|
|
126
|
+
already closed on top of it. The ledger exists so that promise is a row someone has to
|
|
127
|
+
close, not a sentence someone has to remember.
|
|
128
|
+
|
|
129
|
+
### The reserved-to-Auditor steps list (in the execution guide)
|
|
130
|
+
|
|
131
|
+
A running list, next to the rules of engagement, not a per-task Observations note:
|
|
132
|
+
|
|
133
|
+
```
|
|
134
|
+
Steps reserved to the Auditor (never the Executor's model tier, whatever it is):
|
|
135
|
+
- The credential-adjacent check after a ship task (blocked for every model tier so far).
|
|
136
|
+
- Any commit touching a shared or generated resource (harness blocks it for every tier).
|
|
137
|
+
```
|
|
138
|
+
|
|
139
|
+
The first time a harness or permission block turns out predictable — same step, blocked
|
|
140
|
+
for every Executor tried — add it here once. Every subsequent handoff whose task includes
|
|
141
|
+
that step cites this list instead of the step being rediscovered as a fresh `BLOCKED`
|
|
142
|
+
each time. A note written into one task's Observations after the first occurrence is read
|
|
143
|
+
once and never again; this list is read on every handoff.
|
|
144
|
+
|
|
145
|
+
## Rules of engagement (the Executor follows these)
|
|
146
|
+
|
|
147
|
+
Put these at the head of the execution guide, ordered by how often they get broken.
|
|
148
|
+
Items 1-7 are fixed — copy them as they are. **Item 8 is not written from memory or
|
|
149
|
+
inherited from a different run; it is the first concrete step of starting a new run,
|
|
150
|
+
done before task 1 is drafted:**
|
|
151
|
+
|
|
152
|
+
- Check whether this codebase or team already states these somewhere authoritative — a
|
|
153
|
+
CONTRIBUTING file, CI config, an existing style guide, a prior run's execution guide.
|
|
154
|
+
If so, cite it; do not restate it.
|
|
155
|
+
- If nothing authoritative exists, **look at the project before asking** — grep for
|
|
156
|
+
what's already there rather than handing the user a blank checklist:
|
|
157
|
+
|
|
158
|
+
| Rule | Look for |
|
|
159
|
+
|---|---|
|
|
160
|
+
| Version bump per change | A version field (`package.json`, `pyproject.toml`, a `VERSION` file) and any script that already bumps it |
|
|
161
|
+
| No AI co-author trailers | `git log --grep="Co-Authored-By"` — existing trailers (or their consistent absence) |
|
|
162
|
+
| Fetch-before-push | Multiple contributors or branches in `git log`/`git branch -a` — a solo, single-branch repo needs this less |
|
|
163
|
+
| Quality gate green before push | CI config (`.github/workflows/`, `.gitlab-ci.yml`) or a `test`/`lint`/`build` script already wired together |
|
|
164
|
+
| i18n/locale parity | A locales/translations directory with more than one language file |
|
|
165
|
+
| Dependency lockstep | A lockfile and how strict its existing commit history is about touching it |
|
|
166
|
+
|
|
167
|
+
**Then ask with the finding attached, not a blank menu** — turn each hit into a
|
|
168
|
+
yes/no the user can answer in one word instead of a checklist they have to reason
|
|
169
|
+
through from nothing: *"Found `package.json` version `0.2.1` and no bump script — want
|
|
170
|
+
every task to bump it, and if so, how (semver rule, or you decide per task)?"* /
|
|
171
|
+
*"No CI config found — is there a command that should gate every push, or does this
|
|
172
|
+
repo not have one yet?"* This matters most for a user who isn't deep in the codebase
|
|
173
|
+
themselves (vibe-coding a project, not maintaining one they know by heart) — inspecting
|
|
174
|
+
first is the difference between a real choice and homework.
|
|
175
|
+
A rule with nothing detected for it is still offered, just without a finding attached.
|
|
176
|
+
- **Custom rules stay first-class** — a rule is usable as item 8 once it names what it
|
|
177
|
+
enforces, how it's checked, and what triggers it, whether or not it came from the
|
|
178
|
+
table above. Example: *"No schema migration touches a table over 10k rows without a
|
|
179
|
+
stated rollback plan in the task text, and the migration is run once with `--dry-run`
|
|
180
|
+
before it lands for real."* Names the trigger (a migration over the row threshold),
|
|
181
|
+
the requirement (a rollback plan, in the task text), and the check (`--dry-run` first).
|
|
182
|
+
- Settle it once, at the start of the run. Item 8 does not change task to task.
|
|
183
|
+
|
|
184
|
+
1. **One task at a time, in order.** No batching. No starting a phase whose predecessor
|
|
185
|
+
is not `APPROVED`.
|
|
186
|
+
2. **Never invent a value the codebase, infra, or environment declares.** Ports, origins,
|
|
187
|
+
domains, permission keys, versions, enum members — read them. Needing to ask for one
|
|
188
|
+
means a file was skipped.
|
|
189
|
+
3. **Never guess a cause. Observe it.** Read the log, run the query, print the value.
|
|
190
|
+
When investigating a claim about the code, prefer whatever fast navigation tool this
|
|
191
|
+
project already has over blind grep — a code graph, ctags, an LSP, anything that
|
|
192
|
+
answers "who else reads this" faster than a full-text search. Never assume one is
|
|
193
|
+
installed or working; grep is the guaranteed fallback, not a default to reach past a
|
|
194
|
+
faster tool for. This check is quiet, not a topic — look, then use whatever's there.
|
|
195
|
+
Finding nothing is not itself worth telling the user; only bring it up if a tool was
|
|
196
|
+
found and is why the next command looks unusual.
|
|
197
|
+
4. **Verification is running the thing, not reading the code.** "I reviewed it and it
|
|
198
|
+
looks correct" is reported as `FAILED`.
|
|
199
|
+
5. **Scope is the task text.** Note unrelated problems in *Observations*; do not fix
|
|
200
|
+
them.
|
|
201
|
+
6. **Stop and ask** when the instruction contradicts the code, when a task needs a
|
|
202
|
+
decision the document does not make, or when a previously passing check starts
|
|
203
|
+
failing for reasons unrelated to the change. An empty "Blocked" section across a
|
|
204
|
+
whole run reads as ambiguities resolved silently, not as ambiguities that never
|
|
205
|
+
existed. If you resolve one without stopping, log the decision.
|
|
206
|
+
7. **When the system cannot do what a task asks, that is the deliverable.** Report it
|
|
207
|
+
and stop. A task that ends in a well-argued finding is a success.
|
|
208
|
+
8. Project-specific rules go here: version bumping, commit attribution, dependency
|
|
209
|
+
sync before pushing, i18n parity, design-system checks, whatever this codebase
|
|
210
|
+
already enforces elsewhere. Cite the source file instead of restating it — one
|
|
211
|
+
source of truth. **Settled above, before task 1 — not filled in retroactively.**
|
|
212
|
+
|
|
213
|
+
## Before expanding a phase into tasks
|
|
214
|
+
|
|
215
|
+
Read the deferred items ledger (above) for rows whose "Deferred to" matches the phase
|
|
216
|
+
about to be drafted. Every open row becomes a task, or gets re-deferred by editing the
|
|
217
|
+
row — never dropped silently. This replaces re-deriving the phase's obligations from
|
|
218
|
+
memory or a fresh grep of the whole log; the ledger is what makes that unnecessary.
|
|
219
|
+
|
|
220
|
+
## Reporting back (Executor)
|
|
221
|
+
|
|
222
|
+
A compliance log entry is not addressed to anyone — it is a record. When a task, or the
|
|
223
|
+
last task of a phase, is done, close with a message back to the Auditor, same reasoning
|
|
224
|
+
as the handoff that started the work: whoever picks up the audit may be a fresh session
|
|
225
|
+
too, with nothing but the log to go on unless this exists.
|
|
226
|
+
|
|
227
|
+
```
|
|
228
|
+
ROLE: Executor, reporting on <TASK-ID or "Phase <n>, tasks <first>-<last>">.
|
|
229
|
+
|
|
230
|
+
STATUS: <DONE | BLOCKED | FAILED> — logged in <COMPLIANCE_LOG>, section(s) already
|
|
231
|
+
filled in with literal command output.
|
|
232
|
+
|
|
233
|
+
Verify results: <the one-line summary an Auditor would want before deciding whether to
|
|
234
|
+
re-run everything themselves — not a substitute for that re-run>.
|
|
235
|
+
|
|
236
|
+
Deviations from the task text: <any, with the reasoning — or "none">.
|
|
237
|
+
|
|
238
|
+
Stops logged: <any "Blocked" entries raised mid-task per rule 6 — or "none">.
|
|
239
|
+
|
|
240
|
+
Ready for: <"Audit of Phase <n>" | "the next task, <ID>, once this is reviewed">.
|
|
241
|
+
```
|
|
242
|
+
|
|
243
|
+
Do not narrate confidence ("this should be solid now") in place of the verify results —
|
|
244
|
+
the Auditor is about to re-run everything regardless (rule 4 applies to them too); a
|
|
245
|
+
report's job is to point at the evidence, not to argue for a verdict.
|
|
246
|
+
|
|
247
|
+
## Auditing (Auditor)
|
|
248
|
+
|
|
249
|
+
**Re-run. Do not read the report and agree with it.**
|
|
250
|
+
|
|
251
|
+
1. Run every suite yourself, with the project's own commands. A wrong invocation
|
|
252
|
+
produces a false failure and destroys your credibility for the rest of the audit —
|
|
253
|
+
confirm you're running from the right directory, against the right environment,
|
|
254
|
+
before treating an error as a finding.
|
|
255
|
+
2. Re-execute the negative controls from scratch. Back the file up before mutating it
|
|
256
|
+
and restore from that copy — never with a source-control command that could discard
|
|
257
|
+
other uncommitted work in the same tree.
|
|
258
|
+
3. Check artifacts exist: migrations, screenshots, specs, generated bundles.
|
|
259
|
+
4. Read the checks, not just their names. Ask what ordering or fixture would make a
|
|
260
|
+
passing check pass for the wrong reason.
|
|
261
|
+
5. **Write a probe when a claim is load-bearing.** A temporary check that asserts the
|
|
262
|
+
opposite of what the delivered one asserts tells you in thirty seconds whether the
|
|
263
|
+
behavior is real. Delete it afterward.
|
|
264
|
+
6. Check timestamps before concluding something is missing. Work may have landed after
|
|
265
|
+
you looked.
|
|
266
|
+
7. **A delivered negative control only proves the failure mode it was built to catch.**
|
|
267
|
+
Before signing a task off, ask what other claim in the same delivery is load-bearing
|
|
268
|
+
and has no probe of its own — do not accept the one control that exists as coverage
|
|
269
|
+
for the whole feature.
|
|
270
|
+
8. **Before sending the verdict, check what it's missing, not just what it says.** This
|
|
271
|
+
applies to every verdict, not only a full-phase `APPROVED` or a `CONDITIONAL`/
|
|
272
|
+
`REJECTED` with an obvious remediation — a plain `PASS` on one task, mid-phase, with
|
|
273
|
+
something else already outstanding (an earlier remediation, the next task) needs the
|
|
274
|
+
same discipline: it does not end the message on its own. If anything is left to do —
|
|
275
|
+
a remediation handoff, the next task's handoff, or a re-statement of a handoff already
|
|
276
|
+
sent but not yet acted on — it ships in the same message as the verdict, as the
|
|
277
|
+
literal pasteable block from `references/handoffs.md`, not a sentence describing that
|
|
278
|
+
it's still pending. A verdict is not the deliverable; see "Remediation handoff" in
|
|
279
|
+
`references/handoffs.md`. In Autonomous mode, "ships in the same message" means the
|
|
280
|
+
Auditor launches that subagent in the same turn.
|
|
281
|
+
|
|
282
|
+
### Verdicts
|
|
283
|
+
|
|
284
|
+
| Verdict | Meaning |
|
|
285
|
+
|---|---|
|
|
286
|
+
| `APPROVED` | Every gate independently verified |
|
|
287
|
+
| `CONDITIONAL` | Accepted except for named items; state which and what closes them |
|
|
288
|
+
| `REJECTED` | The deliverable does not do what it claims, even if it is green |
|
|
289
|
+
|
|
290
|
+
A passing check that documents wrong behavior as correct is a **rejection**, not an
|
|
291
|
+
observation. It will defend the defect against whoever tries to fix it later.
|
|
292
|
+
|
|
293
|
+
Separate **process findings** from **code findings**. Work can be accepted on its
|
|
294
|
+
merits while the control that should have caught a defect is recorded as failed. Good
|
|
295
|
+
outcomes do not validate a broken process backward. A condition holds as written or it
|
|
296
|
+
stays open — do not fold an unmet condition into "recorded, not blocking" just to avoid
|
|
297
|
+
holding up a phase; that turns a defect into paperwork. Process findings include: the
|
|
298
|
+
wrong model tier (or an inherited one), a small-model task that improvised instead of
|
|
299
|
+
escalating, a negative control restored with a command that discarded other uncommitted
|
|
300
|
+
work, and a question sent to the owner that belonged to the Auditor.
|
|
301
|
+
|
|
302
|
+
### Closing the run
|
|
303
|
+
|
|
304
|
+
A verdict says whether one phase or task passed. It does not say whether the *run* is
|
|
305
|
+
over — those are different claims, and leaving the second one to tone or inference
|
|
306
|
+
means the user ends up asking "so is that everything?" directly, which is the signal
|
|
307
|
+
this section exists to make unnecessary.
|
|
308
|
+
|
|
309
|
+
Before saying a run is over, check the plan of record itself for a phase that exists on
|
|
310
|
+
paper but was never turned into tasks — not just for a next task in the current
|
|
311
|
+
execution guide. A phase can be planned and never scheduled; "nothing left in the guide
|
|
312
|
+
I'm looking at" and "nothing left in the plan" are not the same check.
|
|
313
|
+
|
|
314
|
+
When every phase is `APPROVED` and the plan of record declares no further phase, say so
|
|
315
|
+
as its own line, not folded into other prose:
|
|
316
|
+
|
|
317
|
+
```
|
|
318
|
+
RUN COMPLETE — <plan of record name>. Every phase APPROVED. No further phase declared.
|
|
319
|
+
Nothing outstanding.
|
|
320
|
+
```
|
|
321
|
+
|
|
322
|
+
Anything short of that — one phase closing, a conditional approval, work still queued —
|
|
323
|
+
uses the verdict-plus-handoff shapes in `references/handoffs.md` instead. `RUN COMPLETE` is reserved for the
|
|
324
|
+
one message that actually ends the need for another Auditor turn on this plan of record.
|
|
325
|
+
|
|
326
|
+
## The Auditor is bound by rule 4 too
|
|
327
|
+
|
|
328
|
+
An order can be wrong in the same way a delivery can be wrong, and it is more
|
|
329
|
+
dangerous when it is: the Executor is told a decided item is a stop, not a choice, so a
|
|
330
|
+
wrong decided item arrives armored against the one person positioned to catch it.
|
|
331
|
+
|
|
332
|
+
The failure has one shape: the Auditor asserted a fact about the system — a field
|
|
333
|
+
exists, a function behaves a certain way, a value has a certain sign — from reasoning
|
|
334
|
+
or a single grep hit, without running the command that would actually confirm it.
|
|
335
|
+
**Deliveries get audited by running them; orders too often get written by reasoning
|
|
336
|
+
about them.** That asymmetry is the whole problem.
|
|
337
|
+
|
|
338
|
+
Two rules close it:
|
|
339
|
+
|
|
340
|
+
**1. An order may not assert what it has not run.** Any claim in an order about a
|
|
341
|
+
field, a function's behavior, or a runtime value carries the command that established
|
|
342
|
+
it, pasted, in the order. Rule 4 above — *verification is running the thing, not
|
|
343
|
+
reading the code* — is not only the Executor's rule. It is the protocol's. If the
|
|
344
|
+
Auditor cannot paste the command, the Auditor cannot assert the fact; it goes in the
|
|
345
|
+
order as a question for the Executor to establish, not as a decision.
|
|
346
|
+
|
|
347
|
+
For anything touching money, permissions, or schema, the cheapest form of this is a
|
|
348
|
+
falsifying probe: write the smallest script that would prove the claim wrong and run
|
|
349
|
+
it before the order goes out.
|
|
350
|
+
|
|
351
|
+
**2. A decided item is not reopenable; a factual claim always is.**
|
|
352
|
+
|
|
353
|
+
| | Reopenable by the Executor? |
|
|
354
|
+
|---|---|
|
|
355
|
+
| A decision — scope, trade-off, design | **No.** Believing otherwise is a stop. |
|
|
356
|
+
| A factual claim — this field exists, this function does X, this value is Y | **Always**, with a pasted command that contradicts it. The evidence outranks the order. |
|
|
357
|
+
|
|
358
|
+
An Executor who runs a command that falsifies an order is not exceeding scope; that is
|
|
359
|
+
the second pair of eyes the arrangement exists to buy.
|
|
360
|
+
|
|
361
|
+
## Maintaining the log
|
|
362
|
+
|
|
363
|
+
A compliance log that only ever gets appended to accumulates duplicate paragraphs —
|
|
364
|
+
the same open condition re-stated verbatim across several entries because it was
|
|
365
|
+
easier to copy the last summary than to write a new one. That is a cost, not a
|
|
366
|
+
neutral habit: it makes the current state of the run more expensive to find for
|
|
367
|
+
whoever reads it next, human or agent. Compact the log periodically — collapse a
|
|
368
|
+
condition that has appeared unchanged across several entries into one current
|
|
369
|
+
statement with a pointer to when it opened. `auditkit lint` flags near-duplicate
|
|
370
|
+
paragraphs so this doesn't have to be caught by eye.
|
|
371
|
+
|
|
372
|
+
**Annex count is a health signal, not just a history.** A phase that accumulates many
|
|
373
|
+
corrections after work started is a phase whose plan was under-verified before work
|
|
374
|
+
started, more often than it is a phase that hit genuine surprises. If one phase is
|
|
375
|
+
generating annexes faster than the others, that is worth naming as a finding about the
|
|
376
|
+
planning step, not only fixing task by task. `auditkit lint` warns past a configurable
|
|
377
|
+
threshold.
|
|
378
|
+
|
|
379
|
+
## Repo conventions this rides on
|
|
380
|
+
|
|
381
|
+
Cite whatever this codebase already enforces — version bumping, commit attribution,
|
|
382
|
+
fetch-before-push, environment-value discipline, evidence-driven debugging — from its
|
|
383
|
+
own source of truth rather than restating it here, so there's one place it can drift
|
|
384
|
+
out of sync from.
|
|
385
|
+
|
|
386
|
+
## The CLI
|
|
387
|
+
|
|
388
|
+
If `auditkit` is installed, prefer it over doing these by hand:
|
|
389
|
+
|
|
390
|
+
- `auditkit init <dir>` — scaffold the plan of record, execution guide, and compliance
|
|
391
|
+
log from templates, plus an empty `annexes/` directory.
|
|
392
|
+
- `auditkit lint <dir>` — cross-check task IDs between the execution guide and the
|
|
393
|
+
compliance log in both directions, flag `DONE` reports with no pasted verify output
|
|
394
|
+
(a lone "n/a" or "TBD" is not output),
|
|
395
|
+
flag near-duplicate paragraphs, warn on annex count per phase, flag gates with no
|
|
396
|
+
stated negative control (it must name the action that removes the protection and the
|
|
397
|
+
failure that follows; a negated, deferred, or waived mention such as "n/a" or "TBD" does
|
|
398
|
+
not count). Missing documents, an execution guide with no tasks, and tasks without a
|
|
399
|
+
`**Report:**` line are errors, not a clean result. Lint checks that evidence is
|
|
400
|
+
present, not that it is authentic: re-running it is still the Auditor's job.
|
|
401
|
+
- `auditkit negcontrol --file <path> --break-cmd "<cmd>" --test-cmd "<cmd>"` — backs the
|
|
402
|
+
file up, runs the break command, runs the test (expects failure), restores from the
|
|
403
|
+
backup, runs the test again (expects success), and prints a paste-ready transcript.
|
|
404
|
+
With `--restore-cmd`, the file must still end byte-identical to the backup or it is
|
|
405
|
+
restored from it; `--timeout` bounds each command.
|
|
406
|
+
- `auditkit status <dir>` — combines the status board verdicts with the latest report
|
|
407
|
+
per ID and prints what's still open. Only `APPROVED` closes an item; a `DONE` with no
|
|
408
|
+
verdict is listed as awaiting audit.
|
|
409
|
+
|
|
410
|
+
See the repo's `README.md` for install instructions.
|