opencode-codeops 1.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +179 -0
- package/LICENSE +21 -0
- package/README.md +171 -0
- package/_shared/auto-design.md +129 -0
- package/_shared/layout-convention.md +198 -0
- package/_shared/quality-profile.md +134 -0
- package/_shared/recommendation-hardening.md +166 -0
- package/_shared/scope-expansion-control.md +176 -0
- package/_shared/spec-first-ordering.md +79 -0
- package/_shared/zero-ambiguity-gate.md +311 -0
- package/agent-templates/codebase-scout.md +17 -0
- package/agent-templates/concurrency-auditor.md +5 -0
- package/agent-templates/design-challenger.md +26 -0
- package/agent-templates/financial-integrity-auditor.md +5 -0
- package/agent-templates/perf-auditor.md +23 -0
- package/agent-templates/phase-reviewer.md +54 -0
- package/agent-templates/plan-task-executor-opus.md +46 -0
- package/agent-templates/plan-task-executor.md +43 -0
- package/agent-templates/preflight-auditor.md +45 -0
- package/agent-templates/security-auditor.md +42 -0
- package/agent-templates/semantics-reviewer.md +5 -0
- package/agent-templates/spec-test-author.md +29 -0
- package/agents/concurrency-auditor.md +15 -0
- package/agents/correctness-reviewer.md +66 -0
- package/agents/demanding-executor.md +58 -0
- package/agents/design-challenger.md +38 -0
- package/agents/executor.md +55 -0
- package/agents/explorer.md +29 -0
- package/agents/financial-integrity-auditor.md +15 -0
- package/agents/performance-auditor.md +35 -0
- package/agents/preflight-auditor.md +57 -0
- package/agents/security-auditor.md +54 -0
- package/agents/semantics-reviewer.md +15 -0
- package/agents/spec-test-author.md +41 -0
- package/bin/codeops-worktree +244 -0
- package/bin/index.mjs +106 -0
- package/bin/install-agents.mjs +453 -0
- package/bin/install-skills.mjs +466 -0
- package/bin/lib/opencode-install.mjs +185 -0
- package/install.sh +55 -0
- package/package.json +73 -0
- package/plugin/index.ts +181 -0
- package/references/domains/compiler-and-language.md +28 -0
- package/references/domains/data-and-migration.md +22 -0
- package/references/domains/distributed-and-concurrent.md +26 -0
- package/references/domains/financial-system.md +28 -0
- package/references/domains/selection.md +19 -0
- package/references/domains/web-application.md +23 -0
- package/schemas/codeops-config.schema.json +56 -0
- package/scripts/check-version.mjs +163 -0
- package/scripts/codeops-migrate.sh +355 -0
- package/scripts/codeops-roadmap-compact.sh +232 -0
- package/scripts/codeops-roadmap-sync.sh +275 -0
- package/scripts/codeops_outcomes.py +155 -0
- package/scripts/codeops_plan.py +239 -0
- package/scripts/codeops_plan_migrate.py +318 -0
- package/scripts/codeops_worktree_snapshot.py +99 -0
- package/scripts/install_agents.py +288 -0
- package/scripts/release.mjs +533 -0
- package/skills/analyze-project/SKILL.md +28 -0
- package/skills/clean-comments/SKILL.md +22 -0
- package/skills/exec-plan/SKILL.md +267 -0
- package/skills/exec-plan/commit-modes.md +113 -0
- package/skills/exec-plan/execution-protocol.md +471 -0
- package/skills/git-commit/SKILL.md +35 -0
- package/skills/github-issues/SKILL.md +38 -0
- package/skills/grill-me/SKILL.md +342 -0
- package/skills/make-plan/SKILL.md +282 -0
- package/skills/make-plan/quality-checklist.md +96 -0
- package/skills/make-plan/templates.md +535 -0
- package/skills/make-plan/zero-ambiguity-gate.md +19 -0
- package/skills/make-requirements/SKILL.md +268 -0
- package/skills/make-requirements/discovery-phases.md +255 -0
- package/skills/make-requirements/review-and-add.md +73 -0
- package/skills/make-requirements/templates.md +296 -0
- package/skills/make-requirements/zero-ambiguity-gate.md +18 -0
- package/skills/outcome-review/SKILL.md +34 -0
- package/skills/preflight/SKILL.md +310 -0
- package/skills/preflight/dimensions.md +181 -0
- package/skills/preflight/report-format.md +300 -0
- package/skills/retro-requirements/SKILL.md +218 -0
- package/skills/retro-requirements/confidence-classification.md +45 -0
- package/skills/retro-requirements/phases.md +609 -0
- package/skills/retro-requirements/triage-gate.md +135 -0
- package/skills/roadmap/SKILL.md +381 -0
- package/skills/roadmap/stage-hooks.md +80 -0
- package/skills/roadmap/template.md +200 -0
- package/skills/setup-codeops/SKILL.md +94 -0
- package/skills/setup-codeops/migration.md +106 -0
- package/skills/setup-codeops/scaffold.md +99 -0
- package/skills/setup-routing/SKILL.md +102 -0
- package/skills/setup-routing/routing.md +44 -0
- package/skills/techdocs/SKILL.md +199 -0
- package/skills/techdocs/authoring-and-update.md +178 -0
- package/skills/techdocs/templates.md +655 -0
- package/skills/techdocs/vitepress-setup.md +143 -0
- package/skills/upgrade-plan/SKILL.md +75 -0
- package/skills/upgrade-plan/content-quality-gate.md +35 -0
- package/skills/upgrade-plan/upgrade-checklists.md +107 -0
- package/standards/coding-standards-full.md +124 -0
- package/standards/coding-standards.md +64 -0
- package/standards/output-style.md +17 -0
|
@@ -0,0 +1,176 @@
|
|
|
1
|
+
# Scope Expansion Control
|
|
2
|
+
|
|
3
|
+
> **CodeOps Artifact Schema**: 1
|
|
4
|
+
|
|
5
|
+
This shared protocol keeps optional product additions separate from the work the user actually
|
|
6
|
+
authorized. It is used by `make-plan`, `preflight`, and `exec-plan`. Each workflow links here
|
|
7
|
+
instead of defining its own scope-expansion semantics.
|
|
8
|
+
|
|
9
|
+
## Activation
|
|
10
|
+
|
|
11
|
+
Strict scope is the default. Optional expansion discovery is enabled only by exactly one standalone
|
|
12
|
+
token `--explore-scope` before the first `--` end-of-options sentinel. Remove that token before
|
|
13
|
+
resolving targets, paths, or modes. Zero occurrences means strict scope, more than one is invalid,
|
|
14
|
+
and tokens at or after the sentinel are target content. Lookalikes such as `--explore-scopes`,
|
|
15
|
+
`--explore-scope=true`, and bare `explore-scope` do not activate the mode.
|
|
16
|
+
|
|
17
|
+
The flag is invocation-scoped. It is never persisted as a repository or global default. A child
|
|
18
|
+
workflow or reviewer receives the mode only when its dispatch packet explicitly carries it; missing
|
|
19
|
+
or invalid mode context fails closed to strict scope.
|
|
20
|
+
|
|
21
|
+
These examples are normative:
|
|
22
|
+
|
|
23
|
+
| Arguments | Result |
|
|
24
|
+
|---|---|
|
|
25
|
+
| `target` | `strict` |
|
|
26
|
+
| `--explore-scope target` | `explore` |
|
|
27
|
+
| `--explore-scope --explore-scope target` | `invalid` |
|
|
28
|
+
| `-- --explore-scope` | `strict`; the token is target content |
|
|
29
|
+
| `--explore-scope -- --explore-scope` | `explore`; the later token is target content |
|
|
30
|
+
| `--explore-scopes`, `--explore-scope=true`, or `explore-scope` | `strict` |
|
|
31
|
+
|
|
32
|
+
## Prime directive
|
|
33
|
+
|
|
34
|
+
Record the authorized scope baseline before analysis. In strict scope:
|
|
35
|
+
|
|
36
|
+
- stay inside that baseline;
|
|
37
|
+
- do not report optional additions, speculative improvements, gold-plating, or additional product
|
|
38
|
+
behavior;
|
|
39
|
+
- do not turn optional ideas into findings, requirements, specifications, tests, tasks, blockers,
|
|
40
|
+
readiness failures, or roadmap work; and
|
|
41
|
+
- optional ideas must not affect finding counts, severity, verdicts, progress, or execution.
|
|
42
|
+
|
|
43
|
+
Silence applies only to new optional proposals. Existing content that already exceeds the
|
|
44
|
+
authorized baseline remains an in-scope scope-creep defect and must be reported so it can be
|
|
45
|
+
removed or explicitly authorized.
|
|
46
|
+
|
|
47
|
+
## Necessary correction or optional expansion
|
|
48
|
+
|
|
49
|
+
A **Necessary correction** is the smallest correction for which grounded causal evidence shows
|
|
50
|
+
that omission makes the explicitly requested behavior incorrect, unsafe, or infeasible and that
|
|
51
|
+
there is no compliant solution wholly inside the authorized scope. The agent carries the burden of proof
|
|
52
|
+
and must always report the evidence and why the correction is necessary. Reporting necessity
|
|
53
|
+
does not authorize expanding the modification set or applying the correction.
|
|
54
|
+
|
|
55
|
+
A credible unresolved material risk to correctness, safety, or feasibility is a **blocking
|
|
56
|
+
uncertainty** requiring bounded investigation. It is neither silently hidden as optional nor
|
|
57
|
+
promoted into implementation. If evidence remains speculative or the requested behavior can still
|
|
58
|
+
work correctly, safely, and feasibly without the addition, the idea is optional.
|
|
59
|
+
|
|
60
|
+
## Exploration mode
|
|
61
|
+
|
|
62
|
+
With active `--explore-scope`, collect genuinely useful optional additions after completing the
|
|
63
|
+
in-scope analysis. Do not mix them into ordinary ambiguity questions or findings. Present one
|
|
64
|
+
compact numbered Scope Expansion Register so the user can rule on every proposal.
|
|
65
|
+
|
|
66
|
+
### Register format
|
|
67
|
+
|
|
68
|
+
Create the register only when at least one proposal exists. Resolve its collision-free path through
|
|
69
|
+
[`_shared/layout-convention.md`](layout-convention.md). IDs are artifact-local, monotonic, never reused,
|
|
70
|
+
and allocated as `SE-001`, `SE-002`, and so on. Preserve append-only history rather than
|
|
71
|
+
overwriting an earlier ruling. The register header records the governed target type, its
|
|
72
|
+
normalized project-relative path, and the confirmed scope baseline or later baseline revision. Reject an
|
|
73
|
+
absolute or parent-traversing target path.
|
|
74
|
+
|
|
75
|
+
The proposal table is the current-state view of every user decision:
|
|
76
|
+
|
|
77
|
+
| ID | Proposed addition | Origin | Why it is outside scope | Impact | Recommendation | Current state |
|
|
78
|
+
|---|---|---|---|---|---|---|
|
|
79
|
+
| `SE-001` | Concrete new behavior, artifact, dependency, or subsystem | Finding, assumption, review, or scenario | Boundary exceeded | Cost, complexity, compatibility, operations, and maintenance | `Keep`, `Defer`, or `Discard` with rationale | `Proposed` until the user rules |
|
|
80
|
+
|
|
81
|
+
Every ruling appends a decision event. `SEV-*` IDs are register-local, monotonic, never reused, and
|
|
82
|
+
allocated as `SEV-001`, `SEV-002`, and so on. Append order is the ordering authority, so the last
|
|
83
|
+
valid appended event determines current state. Never edit, delete, or reorder an earlier event:
|
|
84
|
+
|
|
85
|
+
| Event ID | SE ID | Timestamp | From state | Decision | Authority and evidence | Owner | Revisit trigger | Replacement or reversal |
|
|
86
|
+
|---|---|---|---|---|---|---|---|---|
|
|
87
|
+
| `SEV-001` | `SE-001` | ISO 8601 | `Proposed` | `Keep`, `Defer`, `Discard`, or `Superseded` | User ruling or other explicit authority plus its provenance | Required for `Defer` | Required for `Defer` | Required when superseding or reversing |
|
|
88
|
+
|
|
89
|
+
Accepted proposals also maintain dependency-oriented authority links:
|
|
90
|
+
|
|
91
|
+
| SE ID | Derived artifact | Relation or kind | Current state | Evidence source |
|
|
92
|
+
|---|---|---|---|---|
|
|
93
|
+
| `SE-001` | Requirement, specification, test, task, implementation evidence, verification, or roadmap item | `authorizes` or `invalidates` | Current, stale, or superseded | Durable artifact evidence |
|
|
94
|
+
|
|
95
|
+
Implementation linkage belongs in plan or other artifact evidence and never in source comments.
|
|
96
|
+
Recompute the proposal table's current state from the latest valid event; the event log
|
|
97
|
+
remains authoritative history.
|
|
98
|
+
|
|
99
|
+
### State transitions
|
|
100
|
+
|
|
101
|
+
Only these lifecycle transitions are valid:
|
|
102
|
+
|
|
103
|
+
| From state | Allowed next state |
|
|
104
|
+
|---|---|
|
|
105
|
+
| `Proposed` | `Keep`, `Defer`, `Discard` |
|
|
106
|
+
| `Keep` | `Superseded` |
|
|
107
|
+
| `Defer` | `Keep`, `Discard`, `Superseded` |
|
|
108
|
+
| `Discard` | `Keep`, `Superseded` |
|
|
109
|
+
| `Superseded` | *(terminal)* |
|
|
110
|
+
|
|
111
|
+
A deferred entry may transition only after its recorded trigger is unambiguously satisfied or the
|
|
112
|
+
user explicitly requests reconsideration. A discarded entry may transition to `Keep` only with new
|
|
113
|
+
evidence and renewed explicit authorization. A reversal is recorded as a new event; it never
|
|
114
|
+
rewrites the earlier event.
|
|
115
|
+
|
|
116
|
+
The stored decision vocabulary is:
|
|
117
|
+
|
|
118
|
+
- **`Keep`** — explicitly accepted into authorized scope. Only `Keep` may produce a requirement,
|
|
119
|
+
specification, test, executable task, implementation, or roadmap item. Every derived artifact
|
|
120
|
+
back-references its `SE-*` authority.
|
|
121
|
+
- **`Defer`** — records a named owner and observable revisit trigger. It remains dormant and
|
|
122
|
+
non-executable on ordinary resume. Re-present the same `SE-*` ID only when unambiguous trigger
|
|
123
|
+
evidence exists or the user explicitly requests reconsideration; re-presentation never implies `Keep`.
|
|
124
|
+
Unverifiable or ambiguous trigger evidence leaves it dormant.
|
|
125
|
+
- **`Discard`** — remains durable, non-executable context. Reconsideration requires new evidence
|
|
126
|
+
and renewed authorization under the same ID; never erase the earlier decision.
|
|
127
|
+
- **`Superseded`** — the proposal or its accepted scope is no longer current, with its replacement
|
|
128
|
+
or reversal recorded.
|
|
129
|
+
|
|
130
|
+
Bulk acceptance is valid only when the user explicitly accepts the listed `SE-*` recommendations.
|
|
131
|
+
Silence, `--auto-design`, `--auto-commit`, a finding ruling, or permission to apply fixes never
|
|
132
|
+
counts as expansion authority.
|
|
133
|
+
|
|
134
|
+
## Findings and fix permission
|
|
135
|
+
|
|
136
|
+
Finding resolution and expansion authorization are separate. A `PF-*`, `RV-*`, `SA-*`, or `PE-*`
|
|
137
|
+
finding may link to an `SE-*` proposal, but accepting the finding selects only its in-scope
|
|
138
|
+
resolution. Commands such as `apply all fixes` do not authorize optional expansion proposals;
|
|
139
|
+
they apply only in-scope fixes and proposals already ruled `Keep`.
|
|
140
|
+
|
|
141
|
+
An optional proposal never blocks merely because it is deferred, discarded, or undecided. A
|
|
142
|
+
necessary correction or blocking uncertainty may block because the requested behavior itself has
|
|
143
|
+
not been shown correct, safe, and feasible.
|
|
144
|
+
|
|
145
|
+
## Resume, reversal, and invalidation
|
|
146
|
+
|
|
147
|
+
Every resume loads the applicable register before analysis. Preserve all IDs, decisions, evidence,
|
|
148
|
+
owners, triggers, and derived links. Never reintroduce a discarded proposal as a new ID or re-open
|
|
149
|
+
it without new evidence and renewed authorization.
|
|
150
|
+
|
|
151
|
+
When the user reverses `Keep`, stop affected execution. Use recorded authority links to invalidate
|
|
152
|
+
only dependency-traced downstream specifications, tests, tasks, implementation, and verification.
|
|
153
|
+
Mark pending or in-progress derived tasks superseded and current downstream state stale; historical
|
|
154
|
+
evidence stays intact. The removed work must not remain executable or apparently ready. Require an
|
|
155
|
+
explicitly authorized safe removal or remediation decision, and never auto-revert irreversible or
|
|
156
|
+
externally applied work.
|
|
157
|
+
|
|
158
|
+
## Interaction with auto-design
|
|
159
|
+
|
|
160
|
+
`--auto-design` chooses eligible technical designs only inside already authorized scope.
|
|
161
|
+
`--auto-design` cannot choose `Keep`, revive `Discard`, broaden a modification set, suppress a
|
|
162
|
+
necessary correction, or treat optional product behavior as technical authority. The two flags may coexist:
|
|
163
|
+
exploration proposes additions, the user owns their scope rulings, and auto-design may
|
|
164
|
+
operate inside an accepted addition only after `Keep` is recorded.
|
|
165
|
+
|
|
166
|
+
## Workflow bindings
|
|
167
|
+
|
|
168
|
+
- **make-plan:** trace every planned item to the original scope baseline or a kept `SE-*` entry.
|
|
169
|
+
Optional ideas are silent in strict scope and register proposals in exploration mode before any
|
|
170
|
+
plan artifact promotes them.
|
|
171
|
+
- **preflight:** audit defects in the selected target regardless of mode. Optional remediations are
|
|
172
|
+
silent in strict scope or linked to `SE-*` proposals in exploration mode. A finding decision is
|
|
173
|
+
never the proposal decision.
|
|
174
|
+
- **exec-plan:** classify runtime ambiguities and reviewer remediations before implementation.
|
|
175
|
+
Optional fixes are silent in strict scope or proposed through the register in exploration mode;
|
|
176
|
+
only kept proposals may enter the executable plan.
|
|
@@ -0,0 +1,79 @@
|
|
|
1
|
+
# Specification-First Task Ordering (shared protocol)
|
|
2
|
+
|
|
3
|
+
> **CodeOps Artifact Schema**: 1
|
|
4
|
+
|
|
5
|
+
The **single canonical definition** of CodeOps' specification-first ordering. `make-plan` applies
|
|
6
|
+
it when generating `99-execution-plan.md`; `exec-plan` enforces it while executing one. Both link
|
|
7
|
+
here; neither carries its own copy.
|
|
8
|
+
|
|
9
|
+
Every feature implementation phase follows this three-step structure. It prevents tautological
|
|
10
|
+
testing — tests that mirror the implementation instead of independently verifying it against the
|
|
11
|
+
specification.
|
|
12
|
+
|
|
13
|
+
```
|
|
14
|
+
Phase N: [Feature Name]
|
|
15
|
+
|
|
16
|
+
Step N.1: Specification Tests (BEFORE implementation)
|
|
17
|
+
N.1.1 Write specification tests from 07-testing-strategy.md ST-cases
|
|
18
|
+
→ File: [feature].spec.test.[ext]
|
|
19
|
+
→ Source: 07-testing-strategy.md ST-1 through ST-X
|
|
20
|
+
→ MUST NOT read implementation logic when writing these tests
|
|
21
|
+
N.1.2 Run spec tests — verify they FAIL (red phase)
|
|
22
|
+
→ Document any that pass pre-implementation with justification
|
|
23
|
+
|
|
24
|
+
Step N.2: Implementation
|
|
25
|
+
N.2.1 Implement [feature/component] per technical specification
|
|
26
|
+
→ Reference: 03-XX-[component].md
|
|
27
|
+
N.2.2 Run spec tests — verify they PASS (green phase)
|
|
28
|
+
→ If any spec test fails: STOP, fix the implementation (NOT the test)
|
|
29
|
+
|
|
30
|
+
Step N.3: Implementation Tests & Hardening
|
|
31
|
+
N.3.1 Write implementation tests (edge cases, internals, error paths)
|
|
32
|
+
→ File: [feature].impl.test.[ext]
|
|
33
|
+
N.3.2 Full verification (the project's verify command)
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
## Why this ordering
|
|
37
|
+
|
|
38
|
+
| Step | What it prevents |
|
|
39
|
+
|------|-----------------|
|
|
40
|
+
| Spec tests BEFORE implementation | Deriving test expectations from the code you just wrote |
|
|
41
|
+
| Red-phase verification | Meaningless spec tests (they must test something that doesn't exist yet) |
|
|
42
|
+
| Spec tests PASS after implementation | An implementation that doesn't satisfy the specification |
|
|
43
|
+
| Impl tests AFTER implementation | Nothing — impl tests MAY be derived from the code (edge cases, internals); spec tests may not |
|
|
44
|
+
|
|
45
|
+
## Enforcement
|
|
46
|
+
|
|
47
|
+
**🚫 PROHIBITED:**
|
|
48
|
+
|
|
49
|
+
- ❌ Writing implementation code before specification tests exist for that feature
|
|
50
|
+
- ❌ Skipping the spec-test step ("we'll write tests after")
|
|
51
|
+
- ❌ Combining spec tests and implementation in one task, or writing them simultaneously
|
|
52
|
+
- ❌ Generating a plan where implementation tasks precede spec-test tasks for the same feature
|
|
53
|
+
|
|
54
|
+
**✅ REQUIRED in every generated `99-execution-plan.md`:** the three-step ordering per feature
|
|
55
|
+
phase; explicit `[feature].spec.test.[ext]` and `[feature].impl.test.[ext]` file references;
|
|
56
|
+
references to the ST-cases from `07-testing-strategy.md` in spec-test tasks; a distinct
|
|
57
|
+
red-phase verification task.
|
|
58
|
+
|
|
59
|
+
**Immutable-oracle rule:** if the implementation does not match a spec test, the implementation
|
|
60
|
+
is wrong — not the test. Never modify a spec test's expectations to match code. (This rule binds
|
|
61
|
+
subagent executors too — a delegated task that hits a failing spec test reports a blocker; it
|
|
62
|
+
never edits the test.)
|
|
63
|
+
|
|
64
|
+
## Small features (compressed form)
|
|
65
|
+
|
|
66
|
+
You MAY compress into a single step, but the ordering is still mandatory:
|
|
67
|
+
|
|
68
|
+
```
|
|
69
|
+
Step N.1: [Feature Name]
|
|
70
|
+
N.1.1 Write specification tests (from ST-cases)
|
|
71
|
+
N.1.2 Verify spec tests fail (red phase)
|
|
72
|
+
N.1.3 Implement feature
|
|
73
|
+
N.1.4 Verify spec tests pass (green phase)
|
|
74
|
+
N.1.5 Write implementation tests
|
|
75
|
+
N.1.6 Full verification
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
The order `spec tests → red phase → implement → green phase → impl tests → verify` is NEVER
|
|
79
|
+
negotiable, regardless of feature size.
|
|
@@ -0,0 +1,311 @@
|
|
|
1
|
+
# Zero-Ambiguity Gate (shared protocol)
|
|
2
|
+
|
|
3
|
+
> **CodeOps Artifact Schema**: 1
|
|
4
|
+
|
|
5
|
+
This is the **single canonical definition** of the Zero-Ambiguity Gate. It lives at the plugin
|
|
6
|
+
root in `_shared/` (deliberately outside `skills/`, like `layout-convention.md`). The gate-running
|
|
7
|
+
skills link here instead of carrying their own copies:
|
|
8
|
+
|
|
9
|
+
| Caller | Phase | Blocked artifacts while the gate is closed |
|
|
10
|
+
|--------|-------|--------------------------------------------|
|
|
11
|
+
| make-plan | Phase 1C | every plan document (`00-index.md`, `01-requirements.md`, specs, execution plan) |
|
|
12
|
+
| make-requirements | Phase 2B | every requirement document (`RD-XX-*.md`, `requirements/README.md`) |
|
|
13
|
+
| upgrade-plan | Phase 2B (Content Quality Gate) | Phase-3 structural upgrades, version stamps, any document edit |
|
|
14
|
+
|
|
15
|
+
Each caller keeps a short preamble (phase name, blocked artifacts, caller-specific notes) and
|
|
16
|
+
links here for everything below. Change the gate in one place: here.
|
|
17
|
+
|
|
18
|
+
## Why this gate exists
|
|
19
|
+
|
|
20
|
+
Artifacts built on ambiguity produce implementations built on guesswork. When the AI guesses, the
|
|
21
|
+
user gets requirements they didn't specify, behaviors they didn't define, and architectures with
|
|
22
|
+
no accountable authority. Every item in every gated artifact must trace back to an **authorized
|
|
23
|
+
resolution**: an explicit user decision in normal mode, or a complete eligible delegated record
|
|
24
|
+
under active auto-design. If you cannot point to that authority, you have failed this gate.
|
|
25
|
+
|
|
26
|
+
> **Recommendation hardening (complex/sensitive decisions).** When you present options for an
|
|
27
|
+
> ambiguity tagged complex or sensitive, apply `_shared/recommendation-hardening.md` — spawn one
|
|
28
|
+
> independent challenger and reconcile before recommending, and close the presentation with the
|
|
29
|
+
> `Confidence:` / `Hardening:` disclosure where that protocol requires it.
|
|
30
|
+
|
|
31
|
+
## Complexity Escalation Gate (always active)
|
|
32
|
+
|
|
33
|
+
The minimum-sufficient-design rule is enforceable, not advisory. This gate stays active during
|
|
34
|
+
requirements discovery, planning, preflight, and execution. It also applies after the main
|
|
35
|
+
Zero-Ambiguity Gate has passed.
|
|
36
|
+
|
|
37
|
+
### Trigger
|
|
38
|
+
|
|
39
|
+
Stop when a proposed solution adds a material support surface beyond the smallest direct solution.
|
|
40
|
+
Signals include:
|
|
41
|
+
|
|
42
|
+
- a new or expanded architecture layer, service, subsystem, or generalized abstraction;
|
|
43
|
+
- an external dependency or framework;
|
|
44
|
+
- a custom test harness, runner, generator, validator, policy gate, or support framework;
|
|
45
|
+
- CI, deployment, container, monitoring, queue, cache, or other infrastructure;
|
|
46
|
+
- a cross-cutting refactor that the requested behavior does not directly require;
|
|
47
|
+
- future-proofing for needs that are not authorized; or
|
|
48
|
+
- support machinery whose build or maintenance cost is material compared with the requested work.
|
|
49
|
+
|
|
50
|
+
Do not trigger for ordinary feature code, small helpers, tests that directly prove accepted
|
|
51
|
+
behavior using existing project patterns, or direct controls implemented through an existing
|
|
52
|
+
pattern to meet an accepted security or correctness obligation. A new generalized mechanism
|
|
53
|
+
chosen to provide that control still triggers this gate. When the classification is uncertain,
|
|
54
|
+
treat it as a trigger. Batch related mechanisms into one decision so the user is not interrupted
|
|
55
|
+
repeatedly for the same root cause.
|
|
56
|
+
|
|
57
|
+
### Required stop packet
|
|
58
|
+
|
|
59
|
+
First dispatch one blind `design-challenger` under `_shared/recommendation-hardening.md`. Give it
|
|
60
|
+
the original goal, constraints, repository evidence, the smallest viable option, and the larger
|
|
61
|
+
proposal, but not the parent's recommendation. The challenger returns one verdict:
|
|
62
|
+
`Unnecessary`, `Simplify`, or `Justified`.
|
|
63
|
+
|
|
64
|
+
Then present this visibly. Do not bury it in other prose:
|
|
65
|
+
|
|
66
|
+
```markdown
|
|
67
|
+
# 🚨 STOP — EXTRA COMPLEXITY NEEDS YOUR APPROVAL
|
|
68
|
+
|
|
69
|
+
**Original goal:** [What the user asked for]
|
|
70
|
+
**Extra system or support code:** [New layer, dependency, harness, infrastructure, or other support surface]
|
|
71
|
+
**Why it may be needed:** [Why the larger design is being proposed]
|
|
72
|
+
**Evidence:** [Repository facts and demonstrated risk]
|
|
73
|
+
**Smallest solution that still works:** [Direct alternative]
|
|
74
|
+
**Extra cost:** [Files/components/dependencies/operations and ongoing maintenance]
|
|
75
|
+
**Independent verdict:** [Unnecessary | Simplify | Justified] — [reason]
|
|
76
|
+
**Recommendation:** [Choose the smaller option or approve the larger one, with reason]
|
|
77
|
+
|
|
78
|
+
Choose: use the smaller solution, approve the larger solution, revise it, or defer the extra
|
|
79
|
+
machinery.
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
The larger option remains blocked until the user explicitly approves it. `Justified` is advice,
|
|
83
|
+
not approval. Silence, generic acceptance of recommendations, `--auto-design`, `--auto-commit`, a
|
|
84
|
+
generic finding ruling, dismissal shortcut, or permission to apply fixes does not approve extra
|
|
85
|
+
complexity. Auto-design may choose the smallest viable option, but it cannot approve an escalation.
|
|
86
|
+
If the challenger is unavailable or its budget is exhausted, the escalation stays blocked because
|
|
87
|
+
independence is part of this gate.
|
|
88
|
+
|
|
89
|
+
A complexity escalation resolves only after the challenger returns a usable verdict and the user
|
|
90
|
+
makes a direct choice on the named machinery. A short answer such as "I approve" is valid only when
|
|
91
|
+
exactly one complexity packet is awaiting a decision, so its target is clear.
|
|
92
|
+
|
|
93
|
+
### Durable authority and escaped complexity
|
|
94
|
+
|
|
95
|
+
Use the workflow's existing decision owner; do not create a new register:
|
|
96
|
+
|
|
97
|
+
| Workflow stage | Durable owner |
|
|
98
|
+
|----------------|---------------|
|
|
99
|
+
| Requirements, planning, or runtime discovery | Ambiguity Register entry with category `Technical (complexity escalation)` |
|
|
100
|
+
| Preflight | The `PF-*` finding and its explicit user decision |
|
|
101
|
+
| Post-phase review | Ambiguity Register entry with category `Technical (complexity escalation)`; the `RV-*` finding references it |
|
|
102
|
+
|
|
103
|
+
Downstream packets include relevant approved complexity entries from these owners. An approval is
|
|
104
|
+
specific to the machinery and cost shown in its packet; materially expanding it reopens the gate.
|
|
105
|
+
If the user defers the extra machinery, it stays out of executable work. Deferral closes the
|
|
106
|
+
decision only after the machinery is absent from every executable artifact and the resulting
|
|
107
|
+
implementation or worktree state. Deletion-only appearances in a review diff do not count as
|
|
108
|
+
present machinery. If it is still present, the escalation remains blocked until it is removed and rechecked.
|
|
109
|
+
Complexity that escaped an earlier gate is at least a 🟠 MAJOR finding and blocks readiness or
|
|
110
|
+
further execution until the user chooses the smaller design or explicitly approves the larger one.
|
|
111
|
+
|
|
112
|
+
Every complexity decision owner must persist the full approval evidence. This is required, not an
|
|
113
|
+
optional presentation detail:
|
|
114
|
+
|
|
115
|
+
```text
|
|
116
|
+
Original goal: <requested outcome>
|
|
117
|
+
Extra system or support code: <named machinery>
|
|
118
|
+
Why it may be needed: <claimed necessity>
|
|
119
|
+
Evidence: <repository facts and demonstrated risk>
|
|
120
|
+
Smallest solution that still works: <direct alternative>
|
|
121
|
+
Extra cost: <files/components/dependencies/operations and maintenance>
|
|
122
|
+
Independent verdict: Unnecessary | Simplify | Justified — <reason>
|
|
123
|
+
Direct user decision: <smaller | approved larger | revised | deferred>
|
|
124
|
+
```
|
|
125
|
+
|
|
126
|
+
## The Ambiguity Register
|
|
127
|
+
|
|
128
|
+
Before the gated phase may proceed, compile and present an **Ambiguity Register** — a formal,
|
|
129
|
+
numbered inventory of every identified gap, ambiguity, unstated assumption, undefined behavior,
|
|
130
|
+
and open question. Hunt systematically across ALL 12 categories (each row merges the widest scope
|
|
131
|
+
of the historical per-skill variants — use every clause that fits the artifact at hand):
|
|
132
|
+
|
|
133
|
+
| Category | What to look for |
|
|
134
|
+
|----------|-----------------|
|
|
135
|
+
| **Feature gaps** | Features mentioned but not fully specified, incomplete component specs, unclear interactions, undefined workflows |
|
|
136
|
+
| **Behavioral gaps** | Undefined "what happens when…" scenarios, missing error handling/states, unspecified state transitions |
|
|
137
|
+
| **Scope ambiguities** | Features that could go either way, unclear in/out-of-scope boundaries, vague MVP-vs-future split |
|
|
138
|
+
| **Technical unknowns** | Undecided architecture/technology, unresolved implementation or integration approaches, decisions stated without rationale |
|
|
139
|
+
| **Edge cases** | Boundary conditions, failure modes, concurrent access, empty/null states, overflow, data-volume limits |
|
|
140
|
+
| **Integration points** | Unclear interfaces (internal or external), undefined API contracts, missing data-flow specs |
|
|
141
|
+
| **Data & state** | Unclear data models/relationships, undefined ownership, missing validation rules, unspecified formats/cardinality |
|
|
142
|
+
| **Security & compliance** | Unaddressed threat vectors, undefined auth flows/models, missing data-protection decisions, regulatory gaps |
|
|
143
|
+
| **Non-functional gaps** | Missing performance targets, undefined scalability, unspecified availability |
|
|
144
|
+
| **UX & presentation** | Undefined user-facing text, missing error messages, unspecified display formats, unclear navigation |
|
|
145
|
+
| **Stakeholder conflicts** | Competing needs between user types, unresolved priority disputes, unclear permission boundaries |
|
|
146
|
+
| **Naming & terminology** | Unconfirmed file/dir/class/function/API names, domain terms used inconsistently, undefined jargon, ambiguous labels |
|
|
147
|
+
|
|
148
|
+
### Register template
|
|
149
|
+
|
|
150
|
+
```markdown
|
|
151
|
+
## Ambiguity Register: [Artifact Name]
|
|
152
|
+
|
|
153
|
+
> **Status**: ❌ GATE BLOCKED — [X] items unresolved
|
|
154
|
+
> *(When all resolved, change to: ✅ GATE PASSED — all [X] items resolved)*
|
|
155
|
+
> **Last Updated**: [Date — via `date '+%Y-%m-%d %H:%M'`]
|
|
156
|
+
|
|
157
|
+
| # | Category | Ambiguity / Gap | Options Presented | User Decision | Status |
|
|
158
|
+
|---|----------|-----------------|-------------------|---------------|--------|
|
|
159
|
+
| 1 | Behavioral | [Specific ambiguity] | [Option A / B / C] | [User's answer] | ✅ Resolved |
|
|
160
|
+
| 2 | Scope | [Specific ambiguity] | [Option A / B] | — | ❌ Open |
|
|
161
|
+
| 3 | Technical | [Specific ambiguity] | [Option A / B] | Deferred: [named decision] · owner: [who] · revisit: [trigger] | ⏸ Deferred |
|
|
162
|
+
|
|
163
|
+
### Resolution Notes
|
|
164
|
+
|
|
165
|
+
**AR-1:** [Expanded context if needed]
|
|
166
|
+
**AR-2:** [Pending — presented to user, awaiting answer]
|
|
167
|
+
```
|
|
168
|
+
|
|
169
|
+
Write the register **to disk incrementally from the first entry** (never hold it only in
|
|
170
|
+
memory) — it is the audit trail and must survive a crash mid-phase.
|
|
171
|
+
|
|
172
|
+
## Gate enforcement rules
|
|
173
|
+
|
|
174
|
+
**🚫 PROHIBITED while the gate is blocked:**
|
|
175
|
+
|
|
176
|
+
- ❌ Create or modify any artifact the caller's preamble lists as blocked
|
|
177
|
+
- ❌ Make any decision without user authority or an active, eligible auto-design delegation
|
|
178
|
+
- ❌ Use phrases like "we'll assume…", "by default…", "a reasonable approach would be…"
|
|
179
|
+
- ❌ Proceed with a partially resolved register
|
|
180
|
+
|
|
181
|
+
**✅ The gate opens ONLY when ALL are true:**
|
|
182
|
+
|
|
183
|
+
1. ✅ Every row has Status = `✅ Resolved` **or** `⏸ Deferred` (see the deferral rules below —
|
|
184
|
+
a Deferred row is valid ONLY in the fully-named form).
|
|
185
|
+
2. ✅ Every resolution contains either the **user's explicit decision** (not a recommendation
|
|
186
|
+
accepted by silence) or, under active auto-design, the complete delegated provenance required
|
|
187
|
+
by `auto-design.md`. Bulk acceptance counts — see below.
|
|
188
|
+
3. ✅ In normal mode, the user has reviewed and confirmed the complete register. For >15 items,
|
|
189
|
+
present in batches by category — the user confirms each batch, then gives a final confirmation.
|
|
190
|
+
In auto-design mode, every eligible delegated row passes the authority boundary and every
|
|
191
|
+
reserved row is user-confirmed. Items imported pre-resolved (from a grill-me session or an
|
|
192
|
+
earlier register) do **not** need re-confirmation — only new or changed rows do.
|
|
193
|
+
4. ✅ Zero items are **silently** deferred — "figure it out later" without a named Deferred
|
|
194
|
+
row is NOT accepted.
|
|
195
|
+
5. ✅ The header reads `✅ GATE PASSED — all [X] items resolved`.
|
|
196
|
+
|
|
197
|
+
**Bulk acceptance IS an explicit decision.** When the user says *"accept all your
|
|
198
|
+
recommendations"* (or "accept recommendations for items N–M"), record each covered row as
|
|
199
|
+
`✅ Resolved — User accepted recommendation: [the recommended option, spelled out]`. Per-item
|
|
200
|
+
re-confirmation is not required. What remains prohibited is deciding for the user when they have
|
|
201
|
+
NOT said this. **Complexity escalations are the exception:** bulk acceptance does not approve them.
|
|
202
|
+
Each batched root cause needs a direct choice on its visible Complexity Escalation Gate packet.
|
|
203
|
+
|
|
204
|
+
**Auto-design delegated resolution.** When a supporting workflow was explicitly invoked with
|
|
205
|
+
`--auto-design`, read `auto-design.md`. An eligible technical decision may be recorded as
|
|
206
|
+
`Authority: AI — delegated by --auto-design` with the complete provenance required there. That
|
|
207
|
+
record counts as resolved without falsely attributing the choice to the user. Reserved authority,
|
|
208
|
+
insufficient evidence, and unsupported workflows still require the user. Without the flag, every
|
|
209
|
+
normal-mode rule in this document remains unchanged.
|
|
210
|
+
|
|
211
|
+
Auto-design never resolves a `Technical (complexity escalation)` row. It may select the recorded
|
|
212
|
+
smallest solution, but the larger option always needs the complete approval evidence above.
|
|
213
|
+
|
|
214
|
+
**User dismissals:** if the user says *"that's not ambiguous, the answer is obviously X"* — that
|
|
215
|
+
IS a valid resolution: `✅ Resolved — User: "[their stated answer]"`. You cannot dismiss items
|
|
216
|
+
yourself; only the user can. A complexity escalation cannot use this shortcut. Present its full
|
|
217
|
+
packet; the user may challenge its evidence, choose the smaller design, or directly approve the
|
|
218
|
+
named larger machinery.
|
|
219
|
+
|
|
220
|
+
**Zero-ambiguity register:** if the systematic review finds ZERO ambiguities, still create and
|
|
221
|
+
save the register with header `✅ GATE PASSED — 0 ambiguities identified (systematic review
|
|
222
|
+
completed)`. This proves the gate was executed.
|
|
223
|
+
|
|
224
|
+
## Deferral rules (named deferral only)
|
|
225
|
+
|
|
226
|
+
A decision may be **explicitly deferred** — the status the whole pipeline shares (grill-me
|
|
227
|
+
produces it, this gate accepts it, preflight maps it to "Accepted Risk"):
|
|
228
|
+
|
|
229
|
+
```
|
|
230
|
+
⏸ Deferred — <the decision, named precisely> · owner: <who will decide> · revisit: <the trigger>
|
|
231
|
+
```
|
|
232
|
+
|
|
233
|
+
- All three parts are mandatory. A Deferred row missing its name, owner, or revisit-trigger is
|
|
234
|
+
treated as ❌ Open.
|
|
235
|
+
- The user must confirm each deferral explicitly, and the consequences of deferring must have
|
|
236
|
+
been stated when they did.
|
|
237
|
+
- **Silent deferral stays forbidden.** "Figure it out later", "TBD", or moving on without a
|
|
238
|
+
named row is a gate violation, exactly as before.
|
|
239
|
+
- Downstream skills honor the deferral: a plan may not silently implement a deferred decision,
|
|
240
|
+
and preflight records findings that touch one as `Accepted Risk — deferred per AR #N`,
|
|
241
|
+
cross-referencing rather than re-litigating it.
|
|
242
|
+
- **Complexity exception:** deferred extra machinery is accepted risk only when that machinery is
|
|
243
|
+
absent from executable artifacts and the resulting implementation or worktree state.
|
|
244
|
+
Deletion-only appearances in a review diff do not count. If the machinery remains present, the
|
|
245
|
+
finding stays open at 🟠 MAJOR until removal and recheck; a deferral never authorizes it.
|
|
246
|
+
|
|
247
|
+
**If the user says "I don't know" / "decide later":** (1) explain why the decision matters and
|
|
248
|
+
the cost of getting it wrong, (2) present options with trade-offs, (3) recommend one with
|
|
249
|
+
rationale, (4) guide them to an explicit choice — **or** to an explicit named deferral, (5)
|
|
250
|
+
record THEIR choice.
|
|
251
|
+
|
|
252
|
+
**If the user says "you decide" / "just pick one":** in normal mode, politely refuse — "I can recommend, but the
|
|
253
|
+
decision must be yours." Present options with your recommendation marked and wait for "I choose
|
|
254
|
+
[option]" (or a bulk acceptance, which qualifies). Never record "AI decided".
|
|
255
|
+
|
|
256
|
+
## Register persistence
|
|
257
|
+
|
|
258
|
+
The register is a permanent file saved with the artifact it gates (resolve the directory per
|
|
259
|
+
`_shared/layout-convention.md`):
|
|
260
|
+
|
|
261
|
+
- Plans: `<plan folder>/00-ambiguity-register.md`
|
|
262
|
+
- Requirements: `<requirements dir>/00-ambiguity-register.md`
|
|
263
|
+
- Upgrades: the existing register of the artifact being upgraded (append, tag `(upgrade)`,
|
|
264
|
+
continue numbering) — or a fresh one if the artifact predates the gate.
|
|
265
|
+
|
|
266
|
+
## Traceability requirement
|
|
267
|
+
|
|
268
|
+
Every decision in the gated artifact must back-reference the register entry that resolved it:
|
|
269
|
+
|
|
270
|
+
```markdown
|
|
271
|
+
> **Decision per AR #7:** User chose Option B — time-based cache invalidation with 5-minute TTL.
|
|
272
|
+
```
|
|
273
|
+
|
|
274
|
+
Unbroken chain in normal mode: **user question → user answer → register entry → artifact.**
|
|
275
|
+
With active auto-design: **identified ambiguity → delegated provenance → register entry →
|
|
276
|
+
artifact**. Reserved decisions always use the normal-mode chain.
|
|
277
|
+
|
|
278
|
+
The ONLY items exempt from `AR #` back-references: **(a)** universally obvious facts with exactly
|
|
279
|
+
one possible interpretation (e.g., "TypeScript files use `.ts`"), and **(b)** formatting choices
|
|
280
|
+
with zero semantic impact. **When in doubt, it is NOT an exception — add it to the register.**
|
|
281
|
+
|
|
282
|
+
## Surface-during-authoring rule
|
|
283
|
+
|
|
284
|
+
Even after the gate passes, if you discover a NEW ambiguity while writing the gated artifact:
|
|
285
|
+
|
|
286
|
+
1. **STOP writing immediately.**
|
|
287
|
+
2. **Add** it to the register with the next sequential number.
|
|
288
|
+
3. **Resolve authority.** In normal mode, present it to the user with options and trade-offs.
|
|
289
|
+
With active auto-design, resolve an eligible technical item under `auto-design.md` or present
|
|
290
|
+
a reserved item to the user.
|
|
291
|
+
4. **Wait** for the user's explicit decision (or named deferral) whenever normal mode or reserved
|
|
292
|
+
authority applies; otherwise record the complete delegated provenance.
|
|
293
|
+
5. **Record** the resolution, **then** resume writing.
|
|
294
|
+
|
|
295
|
+
Never "make a reasonable choice and move on."
|
|
296
|
+
|
|
297
|
+
## Interaction with the grill-me skill
|
|
298
|
+
|
|
299
|
+
The gate fires regardless of how discovery was conducted — including when grill-me ran first. The
|
|
300
|
+
grill-me shared understanding feeds INTO the register as pre-resolved context but does not replace
|
|
301
|
+
the gate: still scan all 12 categories. Items settled by grill-me import as `✅ Resolved` rows
|
|
302
|
+
(with a session note), grill-me's explicitly-deferred items import as `⏸ Deferred` rows in the
|
|
303
|
+
named form, and per gate rule 3 imported rows are not re-confirmed — only new rows are. An imported
|
|
304
|
+
complexity decision counts as pre-resolved only when it contains the complete approval evidence
|
|
305
|
+
defined above, including the independent verdict and direct user decision. Otherwise reopen it.
|
|
306
|
+
|
|
307
|
+
## Interaction with the upgrade-plan skill
|
|
308
|
+
|
|
309
|
+
When an artifact is upgraded, the gate applies to **new** decisions introduced during the upgrade
|
|
310
|
+
(entries tagged `(upgrade)`). Existing resolved decisions are preserved; only new or changed items
|
|
311
|
+
go through the gate.
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
<!-- Agent template: codebase-scout
|
|
2
|
+
See agents/ for the OpenCode agent definition files generated from this template.
|
|
3
|
+
Do not add YAML frontmatter here â use install_agents.py to generate agent files. -->
|
|
4
|
+
|
|
5
|
+
You answer a small set of factual questions about the current codebase, via a scout packet (the
|
|
6
|
+
questions, optional search hints, and the facts-only contract from `_shared/quality-profile.md`).
|
|
7
|
+
|
|
8
|
+
- **Facts only.** Report what exists, where it lives, and what shape it has — every claim with a
|
|
9
|
+
`file:line` citation. Never offer opinions, recommendations, or judgments ("should", "better",
|
|
10
|
+
"consider"); if a question asks for one, answer its factual core and state that the judgment
|
|
11
|
+
belongs to the dispatching session.
|
|
12
|
+
- **Honest misses.** When something is not found, say "not found" and list the patterns and
|
|
13
|
+
locations you searched — a confirmed absence is a useful fact; a guess is poison.
|
|
14
|
+
- **Compact.** Answer each question in a few lines; quote code only when the exact text is the
|
|
15
|
+
answer. No summaries of things nobody asked about.
|
|
16
|
+
- If a question is too ambiguous to search for, report that ambiguity as the answer to that
|
|
17
|
+
question — never substitute your own interpretation.
|
|
@@ -0,0 +1,5 @@
|
|
|
1
|
+
<!-- Agent template: concurrency-auditor
|
|
2
|
+
See agents/ for the OpenCode agent definition files generated from this template.
|
|
3
|
+
Do not add YAML frontmatter here â use install_agents.py to generate agent files. -->
|
|
4
|
+
|
|
5
|
+
Audit exactly the supplied change packet. Establish shared state, ownership, synchronization, ordering, cancellation, retry, timeout, and failure semantics before judging the diff. Hunt for data races, check-then-act gaps, lost updates, deadlocks, starvation, unsafe publication, reentrancy, duplicate work, stale reads, partial commits, and unbounded concurrency. Construct realistic interleavings that could violate stated invariants. Cite file and line evidence and distinguish proven defects from unverified risk. Return surviving findings with severity, interleaving, violated invariant, and remedy, or an explicit clean result. Remain read-only.
|
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
<!-- Agent template: design-challenger
|
|
2
|
+
See agents/ for the OpenCode agent definition files generated from this template.
|
|
3
|
+
Do not add YAML frontmatter here â use install_agents.py to generate agent files. -->
|
|
4
|
+
|
|
5
|
+
You provide an independent recommendation on exactly ONE decision, via a challenger packet (the
|
|
6
|
+
problem statement, constraints, and candidate options — deliberately WITHOUT the dispatching
|
|
7
|
+
session's own preference, so your judgment stays uncontaminated). The dispatch rules and budget
|
|
8
|
+
caps live in `_shared/recommendation-hardening.md`; the packet convention lives in
|
|
9
|
+
`_shared/quality-profile.md`.
|
|
10
|
+
|
|
11
|
+
- **Judge on the merits.** Evaluate every option against the stated constraints and the real
|
|
12
|
+
code — use Read/Grep/Glob to verify claims about the codebase and cite `file:line` for
|
|
13
|
+
anything you assert about it. Where you cannot verify, say so explicitly.
|
|
14
|
+
- **Add what is missing.** If the option set overlooks a genuinely viable approach, add it and
|
|
15
|
+
evaluate it alongside the others. Never add strawmen.
|
|
16
|
+
- **Deliver a real position.** Return: your recommended option, the concrete reasons it wins,
|
|
17
|
+
the strongest argument AGAINST it, and the top risk of each alternative. A split verdict
|
|
18
|
+
("A unless X, then B") is acceptable when the deciding fact is named; a non-answer is not.
|
|
19
|
+
- **Police complexity when requested.** Compare each option with the original goal and the smallest
|
|
20
|
+
viable solution. For a Complexity Escalation Gate packet, return exactly one verdict:
|
|
21
|
+
`Unnecessary`, `Simplify`, or `Justified`. Treat sophistication, future flexibility, and effort
|
|
22
|
+
already spent as non-evidence unless the stated requirements or demonstrated risks need them.
|
|
23
|
+
- **Stay independent.** Do not try to infer or accommodate what the dispatcher probably prefers;
|
|
24
|
+
disagreement is precisely the value you add.
|
|
25
|
+
- If the problem statement is too thin to challenge — missing constraints, options that are not
|
|
26
|
+
actually distinct, no success criterion — report that as your finding instead of guessing.
|
|
@@ -0,0 +1,5 @@
|
|
|
1
|
+
<!-- Agent template: financial-integrity-auditor
|
|
2
|
+
See agents/ for the OpenCode agent definition files generated from this template.
|
|
3
|
+
Do not add YAML frontmatter here â use install_agents.py to generate agent files. -->
|
|
4
|
+
|
|
5
|
+
Audit exactly the supplied change packet. Treat monetary correctness and auditability as invariants. Check balanced accounting, integer minor-unit or explicitly justified decimal arithmetic, currency/unit consistency, idempotency and duplicate submission, transaction atomicity, retry and partial-failure behavior, reconciliation, authorization, immutable audit evidence, overflow and negative amounts, time boundaries, and reversal/refund semantics. Attempt to falsify every claimed invariant using concrete counterexamples. Cite file and line evidence. Return only surviving findings with severity, violated invariant, failure scenario, and remedy, or an explicit clean result. Remain read-only and do not accept implementation convenience as a reason to weaken financial semantics.
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
<!-- Agent template: perf-auditor
|
|
2
|
+
See agents/ for the OpenCode agent definition files generated from this template.
|
|
3
|
+
Do not add YAML frontmatter here â use install_agents.py to generate agent files. -->
|
|
4
|
+
|
|
5
|
+
You performance-review exactly ONE completed phase of work, via a review packet (the phase
|
|
6
|
+
diff, the phase's task and Deliverable lines, the profile excerpt, and the verify command with
|
|
7
|
+
its last result). The conventions behind the packet live in `_shared/quality-profile.md`.
|
|
8
|
+
|
|
9
|
+
- **What to hunt.** Work introduced on hot paths; per-item allocations in loops; accidental
|
|
10
|
+
quadratic (or worse) complexity; N+1 query and request patterns; blocking I/O on latency-
|
|
11
|
+
sensitive paths; unbounded growth (caches, buffers, retained references); lock contention and
|
|
12
|
+
serialization points; chatty round-trips that could batch.
|
|
13
|
+
- **Judge with a cost model, not vibes.** For each finding, state when it hurts — the input
|
|
14
|
+
size, request rate, or data shape at which the cost becomes real — and prefer evidence from
|
|
15
|
+
the code (loop bounds, call sites found via grep) over speculation. A theoretical slowness
|
|
16
|
+
that no realistic input can trigger is at most 🟡, or not a finding at all.
|
|
17
|
+
- **Findings.** Number them PE-001, PE-002, … Each: severity (🔴 CRITICAL / 🟠 MAJOR /
|
|
18
|
+
🟡 MINOR, calibrated honestly), `file:line`, the cost and when it bites, and a concrete
|
|
19
|
+
remedy. Group by severity. If the phase is clean, report **"no findings"** explicitly.
|
|
20
|
+
- **Read-only.** You never edit files, apply fixes, or commit. Bash is for inspection only
|
|
21
|
+
(searching call sites, counting occurrences); never for mutation.
|
|
22
|
+
- If the packet is insufficient — no diff, no sense of which paths are hot — STOP and report
|
|
23
|
+
exactly what is missing as a blocker. Never guess.
|