@sabaiway/agent-workflow-engine 4.0.0 → 4.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,37 @@ All notable changes to the methodology engine. Versions are this **package's** n
4
4
  they are distinct from the **deployment-lineage** stamp written into a project's `docs/ai/`
5
5
  (which tracks the shared `agent-workflow` lineage, head `3.0.0`).
6
6
 
7
+ ## 4.2.0 — a zero governing-spec citation names the adoption state it relies on (AD-123)
8
+
9
+ `references/agent-rules-lens.md`'s **Spec-first** bullet, `references/planning.md`'s *Goal and boundary*
10
+ and `references/specs.md`'s *Governing specs are plural* said a plan may cite ZERO governing specs
11
+ "during adoption" — and nothing defined adoption, so "zero, every plan, forever" read exactly like
12
+ adopting. Each now says the same thing at its own point of use: a ZERO names the state it relies on —
13
+ `not adopted` (no store, or a recorded decline), `adopting` (a store with no live contract) or
14
+ `nothing spec-covered touched` (a store with live contracts) — and a bare zero is never a licence; the
15
+ store's own state is what the kit's `status` and upgrade advisor report.
16
+
17
+ The outgoing lens body is appended to `agent-rules-lens-priors.md` (append-only), so every deployed
18
+ `agent_rules.md` on the previous wording refreshes on the next upgrade instead of reading as a custom
19
+ edit; `test/lens-fragment.test.mjs` computes the outgoing body by swapping the Spec-first line back and
20
+ pins the tokens `adoption state` and `never a licence`.
21
+
22
+ ## 4.1.0 — the queue is a named surface with a checker, not a prose promise
23
+
24
+ `references/planning.md` gains **`## The queue`**: `docs/plans/queue.md` NAMES work and never holds
25
+ the analysis of it — a row is one plain sentence saying what the work is and for whom, then its id,
26
+ then a short body, while measurements and `file:line` citations belong to an ADR or the record the
27
+ row points at. A row that goes terminal is DELETED in the same change that mints its closing
28
+ artifact, and where the queue is gitignored the deleted text is first written to a purge archive,
29
+ because there git history is no tombstone. Frozen work with a stated resume condition is not
30
+ terminal and stays, in its own bucket; order inside a bucket IS priority.
31
+
32
+ The canon names its own rung and the command that runs it, because a prose promise to trim later was
33
+ measured failing — 62 dead rows had accumulated by the time anybody counted.
34
+
35
+ `references/procedures.md` points at the new section by named anchor and stays the terse pointer it
36
+ is meant to be (its own test asserts it stays smaller than the canon it binds to).
37
+
7
38
  ## 4.0.0 — the scenario floor, and five answers the canon now states at its own points of use (AD-117)
8
39
 
9
40
  Slice 4 wrote the layer's first real specs and came back with six questions the canon had left for
package/SKILL.md CHANGED
@@ -3,7 +3,7 @@ name: agent-workflow-engine
3
3
  description: Canonical home of the agent-workflow planning methodology — the capped plan shape (goal and boundary, module ledger, verification), plan lifecycle, queue.md series index, mandatory Cleanup phase, the feature-spec canon (the durable per-feature contract layer with its frozen schema and Out-of-scope discipline), the bounded methodology slot fragment, the orchestration-recipe vocabulary (Solo / Reviewed / Council / Delegated), and the activity-procedures canon (plan-authoring / plan-execution, with typed recipe slots). A published, installable npm package (available:true) that *provides* the methodology text; it mutates nothing. The composition root (agent-workflow-kit) reads this canon LIVE from the installed engine and injects the bounded slots from it — one source of truth, no bundled mirror; `npx @sabaiway/agent-workflow-kit@latest init` installs the engine.
4
4
  disable-model-invocation: true
5
5
  metadata:
6
- version: '4.0.0'
6
+ version: '4.2.0'
7
7
  ---
8
8
 
9
9
  # agent-workflow-engine
package/capability.json CHANGED
@@ -3,7 +3,7 @@
3
3
  "schema": 1,
4
4
  "name": "agent-workflow-engine",
5
5
  "kind": "methodology-engine",
6
- "version": "4.0.0",
6
+ "version": "4.2.0",
7
7
  "available": true,
8
8
  "provides": ["plan"],
9
9
  "roles": {},
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@sabaiway/agent-workflow-engine",
3
- "version": "4.0.0",
3
+ "version": "4.2.0",
4
4
  "description": "Canonical home of the agent-workflow planning methodology — the capped plan shape (goal and boundary, module ledger, verification), plan lifecycle, queue.md series index, and mandatory Cleanup phase, consumed by the kit (composition root). The methodology engine of the agent-workflow family.",
5
5
  "keywords": [
6
6
  "ai-agents",
@@ -194,3 +194,22 @@ Apply these when authoring a plan, reviewing, folding a finding, or editing code
194
194
  - **Recipe fidelity.** Council runs every backend the recipe names, **every round**; silently dropping a ready backend for quota/convenience is a forbidden downgrade — an unavailable backend is a LOUD, stated degrade, never a quiet drop.
195
195
  - **ExitPlanMode ≠ execute.** A harness "approved — start coding" prompt authorizes the PLAN only; this methodology overrides it. Continue into execution only as a DELIBERATE transition after the plan + cold-start prompt exist, never an implicit slide.
196
196
  - **Cost lanes.** Route every step to the **cheapest adequate executor** — L0 deterministic script (the batched gate matrix over `gates.json`, the rotation `--check`s) · L1 cheap subagent (extraction/drafting only; the orchestrator verifies) · L2 subscription bridge · L3 frontier judgment. A step with **no named guardrail does not move down** a lane, and the **red lines never move down** (council review models · real code · ADR/handover/changelog-entry wording · persuasive copy · go/no-go · the approval asks). Own-error repair: salvage recorded state first (L0/L1, batched), never frontier re-derivation. **Prompt economy:** read-only fan-out (research/sweeps/extraction) runs ONLY on restricted-tool vehicles — a full-tool subagent for read-only work is a forbidden lane downgrade (invisible prompt-flood + blast radius), and a subagent is never told to shell out for facts obtainable read-only; the orchestrator's own shell form is ONE plain pipeline per call (a `;`/`&&` chain or env-prefixed invocation never matches a prefix allow rule); a fan-out launcher that gates per call yields to the agent-spawn lane — capability-gated: without restricted-tool vehicles (generic full-tool spawning does not count), read-only research stays in the orchestrator's own context, never a vehicle mandate a host cannot satisfy. Judgment, code, synthesis stay at the frontier lane (a task that genuinely runs/writes keeps a full-tool subagent); honest limit: no deterministic gate classifies a dispatch — canon at the point of use + placed vehicles + the retro loop. **Writer economy:** a stage's repeated WRITER commands batch — evidence declarations ride consecutive plain invocations of ONE allow-listed tool, other stage writers combine via one launcher per stage; never an unbatched writer scatter (each gated write is its own prompt).
197
+
198
+ <!-- prior: 2026-08-27 (AD-123) — before the adoption-state clause in the Spec-first bullet -->
199
+ ### 2.x. Planning, review & process-fidelity invariants
200
+ Apply these when authoring a plan, reviewing, folding a finding, or editing code — the layer read **before any code change**. (Full canon: the project's planning / workflow-methodology + orchestration canon. This section is rendered from that canon and refreshed on upgrade; a custom edit is preserved verbatim, but flagged.)
201
+ - **Fold by code, not prose.** Before folding a code-touching finding into a plan or change, read the cited `file:line` and cite it — a prose fold drifts from the code and seeds the next bug.
202
+ - **Finding scope (plan-execution) — name the invariant BEFORE the edit.** During EXECUTION only — a plan under authoring has no shipped behaviour to call a live defect in, so plan-review carries none of this. Every finding names the invariant its fix would enforce, and where that invariant already lives decides the disposition: already an acceptance criterion of the phase → **fold here**; it would have to be ADDED → ship the **narrow fix** for the found site (red first, then green) and queue ONLY the generalization — a deferral row carries the invariant, the origin `file:line`, the narrow fix, its proof and a residual exposure declared NOT live; no correct narrow fix → **blocking**: the phase does not close, and it is **never queued**. Two bars declared before each round: a finding counts only if it changes a **WRITE/REMOVE decision** or is a false statement in shipped text; a repeat finding in one subarea **routes to SUBTRACTION**, not a fourth patch.
203
+ - **Right altitude.** Pin intent + invariants + acceptance criteria (named tests); leave fine code-mechanics to Execute, where prose cannot diverge from reality.
204
+ - **Spec-first.** A plan names its GOVERNING spec(s) — zero, one or many, one per touched spec-covered slice (the feature spec under `docs/ai/specs/`; page-only coverage governs as an ADOPTION SHIM, with Out of scope + Revision stated inline in the plan). Each cited spec's Out of scope bounds that slice's work and the plan's non-goals restate it per slice — no global union; a cross-spec conflict is resolved by a spec revision BEFORE approval, never by silent precedence. A NEW feature's draft spec exists AT plan review (a `create` row); a change to a governed contract rides the plan as its proposed revision (a `modify` row); approval confirms plan and contract atomically, and the revision lands with the code. Scenario bindings are per scenario: a new scenario is `unbound` until its test lands in the same plan, and a status never regresses for an extension.
205
+ - **No code-mechanics in the plan.** A ledger row carries its path and anchor, and Verification carries the exact commands (the plan-shape canon) — checked syntax: the plan's own Verification runs them against an explicit expected outcome or gate; the only other syntax a plan may carry is a literal fixture/schema fragment a named test copies or validates. Un-run, logic-bearing syntax — control-flow, a regex, a glob, a grammar, an algorithm body, a mini-DSL — never lives in plan prose, however plausible or shell-verified it looks: a fold or draft that wants one is the trigger to write the test instead.
206
+ - **Test-as-spec.** Fold a code-touching finding into a red→green TEST, not a prose paragraph — the gate is the only deterministic checker; a paragraph cannot self-check.
207
+ - **Characterize-first.** Before editing UNCOVERED code, pin its current behavior in a green test, then edit — any unintended change goes red. Never edit what has no checker; first give it one. Keep edits atomic/reversible; prefer SUBTRACTIVE folds.
208
+ - **State table BEFORE the guard — enumerate by PROOF, never by exclusion.** The subtraction rule above fires on a repeat finding, which is a LATE signal: by then the review has paid for each miss. The EARLY signal is structural — a decision whose input has **several independent state dimensions** (is it tracked? do the bytes still match the source? does the neighbouring file exist?). Write the table first, admit the write with **ONE conjunction of proven facts**, and funnel every other cell into a single refusal; the table is then the table-driven test. An exclusion list (`if (bad1) return; if (bad2) return;`) fails **OPEN** on the first state nobody enumerated — and "unreadable" is a state, distinct from "absent". A reviewer cannot save you here: it judges the patch in front of it and can only name the NEXT missing state, one round at a time.
209
+ - **Fold minimally — prose has no checker.** An ephemeral, gitignored plan is prose with no executable checker; fold **minimally, in ONE place** and run a **self-consistency** read across the plan before every re-review — a fold that drifts several prose spots is what turns a 2-round review into churn.
210
+ - **Heavy review at the diff.** Plan-review settles architecture only (≤2 rounds, stop at the pre-existing→fold-induced crossover); the exhaustive per-row review runs against real compiling code + the full suite, where a regression fails a gate immediately. **Backend divergence** (one backend grounded-ships while another keeps revising mechanics) IS that crossover — resolve at altitude, don't exhaust the strictest backend; route an all-mechanics/CI or prose-only artifact to a **thin plan + diff-review**.
211
+ - **Convergence bar.** A review loop is CLEAN only when one round returns **0 blockers + 0 majors** from EVERY backend the recipe names (nits + a ship verdict is the stop). Folding ≠ convergence — re-review after folding.
212
+ - **Per-round emission.** Every review round emits **{round N · finding-origin tally · per-backend verdict}** so the crossover is a computed, visible signal, not a remembered rule.
213
+ - **Recipe fidelity.** Council runs every backend the recipe names, **every round**; silently dropping a ready backend for quota/convenience is a forbidden downgrade — an unavailable backend is a LOUD, stated degrade, never a quiet drop.
214
+ - **ExitPlanMode ≠ execute.** A harness "approved — start coding" prompt authorizes the PLAN only; this methodology overrides it. Continue into execution only as a DELIBERATE transition after the plan + cold-start prompt exist, never an implicit slide.
215
+ - **Cost lanes.** Route every step to the **cheapest adequate executor** — L0 deterministic script (the batched gate matrix over `gates.json`, the rotation `--check`s) · L1 cheap subagent (extraction/drafting only; the orchestrator verifies) · L2 subscription bridge · L3 frontier judgment. A step with **no named guardrail does not move down** a lane, and the **red lines never move down** (council review models · real code · ADR/handover/changelog-entry wording · persuasive copy · go/no-go · the approval asks). Own-error repair: salvage recorded state first (L0/L1, batched), never frontier re-derivation. **Prompt economy:** read-only fan-out (research/sweeps/extraction) runs ONLY on restricted-tool vehicles — a full-tool subagent for read-only work is a forbidden lane downgrade (invisible prompt-flood + blast radius), and a subagent is never told to shell out for facts obtainable read-only; the orchestrator's own shell form is ONE plain pipeline per call (a `;`/`&&` chain or env-prefixed invocation never matches a prefix allow rule); a fan-out launcher that gates per call yields to the agent-spawn lane — capability-gated: without restricted-tool vehicles (generic full-tool spawning does not count), read-only research stays in the orchestrator's own context, never a vehicle mandate a host cannot satisfy. Judgment, code, synthesis stay at the frontier lane (a task that genuinely runs/writes keeps a full-tool subagent); honest limit: no deterministic gate classifies a dispatch — canon at the point of use + placed vehicles + the retro loop. **Writer economy:** a stage's repeated WRITER commands batch — evidence declarations ride consecutive plain invocations of ONE allow-listed tool, other stage writers combine via one launcher per stage; never an unbatched writer scatter (each gated write is its own prompt).
@@ -3,7 +3,7 @@ Apply these when authoring a plan, reviewing, folding a finding, or editing code
3
3
  - **Fold by code, not prose.** Before folding a code-touching finding into a plan or change, read the cited `file:line` and cite it — a prose fold drifts from the code and seeds the next bug.
4
4
  - **Finding scope (plan-execution) — name the invariant BEFORE the edit.** During EXECUTION only — a plan under authoring has no shipped behaviour to call a live defect in, so plan-review carries none of this. Every finding names the invariant its fix would enforce, and where that invariant already lives decides the disposition: already an acceptance criterion of the phase → **fold here**; it would have to be ADDED → ship the **narrow fix** for the found site (red first, then green) and queue ONLY the generalization — a deferral row carries the invariant, the origin `file:line`, the narrow fix, its proof and a residual exposure declared NOT live; no correct narrow fix → **blocking**: the phase does not close, and it is **never queued**. Two bars declared before each round: a finding counts only if it changes a **WRITE/REMOVE decision** or is a false statement in shipped text; a repeat finding in one subarea **routes to SUBTRACTION**, not a fourth patch.
5
5
  - **Right altitude.** Pin intent + invariants + acceptance criteria (named tests); leave fine code-mechanics to Execute, where prose cannot diverge from reality.
6
- - **Spec-first.** A plan names its GOVERNING spec(s) — zero, one or many, one per touched spec-covered slice (the feature spec under `docs/ai/specs/`; page-only coverage governs as an ADOPTION SHIM, with Out of scope + Revision stated inline in the plan). Each cited spec's Out of scope bounds that slice's work and the plan's non-goals restate it per slice — no global union; a cross-spec conflict is resolved by a spec revision BEFORE approval, never by silent precedence. A NEW feature's draft spec exists AT plan review (a `create` row); a change to a governed contract rides the plan as its proposed revision (a `modify` row); approval confirms plan and contract atomically, and the revision lands with the code. Scenario bindings are per scenario: a new scenario is `unbound` until its test lands in the same plan, and a status never regresses for an extension.
6
+ - **Spec-first.** A plan names its GOVERNING spec(s) — zero, one or many, one per touched spec-covered slice (the feature spec under `docs/ai/specs/`; page-only coverage governs as an ADOPTION SHIM, with Out of scope + Revision stated inline in the plan). A ZERO names the adoption state it relies on — `not adopted` (no store) or `adopting` (a store with no live contract), either with a recorded decline, or `nothing spec-covered touched` (a store with live contracts) — a bare zero is never a licence; the store's own state is what `status` and the upgrade advisor report. Each cited spec's Out of scope bounds that slice's work and the plan's non-goals restate it per slice — no global union; a cross-spec conflict is resolved by a spec revision BEFORE approval, never by silent precedence. A NEW feature's draft spec exists AT plan review (a `create` row); a change to a governed contract rides the plan as its proposed revision (a `modify` row); approval confirms plan and contract atomically, and the revision lands with the code. Scenario bindings are per scenario: a new scenario is `unbound` until its test lands in the same plan, and a status never regresses for an extension.
7
7
  - **No code-mechanics in the plan.** A ledger row carries its path and anchor, and Verification carries the exact commands (the plan-shape canon) — checked syntax: the plan's own Verification runs them against an explicit expected outcome or gate; the only other syntax a plan may carry is a literal fixture/schema fragment a named test copies or validates. Un-run, logic-bearing syntax — control-flow, a regex, a glob, a grammar, an algorithm body, a mini-DSL — never lives in plan prose, however plausible or shell-verified it looks: a fold or draft that wants one is the trigger to write the test instead.
8
8
  - **Test-as-spec.** Fold a code-touching finding into a red→green TEST, not a prose paragraph — the gate is the only deterministic checker; a paragraph cannot self-check.
9
9
  - **Characterize-first.** Before editing UNCOVERED code, pin its current behavior in a green test, then edit — any unintended change goes red. Never edit what has no checker; first give it one. Keep edits atomic/reversible; prefer SUBTRACTIVE folds.
@@ -25,7 +25,9 @@ independently verifiable boundaries, never by document size — or it is a SWEEP
25
25
 
26
26
  - **Goal and boundary** (10 lines) — the observable outcome, what behaviour is preserved, explicit
27
27
  non-goals, and the GOVERNING spec(s) ([`specs.md`](specs.md)): zero, one or many — one per touched
28
- spec-covered slice — each cited spec's Out of scope restated as a non-goal for that slice.
28
+ spec-covered slice — a ZERO names the adoption state it relies on (not adopted · adopting — either
29
+ with a recorded decline — · nothing spec-covered touched); a bare zero is never a licence. Each
30
+ cited spec's Out of scope is restated as a non-goal for that slice.
29
31
  - **Module ledger** (60 lines) — the single list of paths, and the plan's execution order.
30
32
  - **Verification** (20 lines) — the acceptance check, plus one command that validates the whole ledger.
31
33
  - **Phase: Cleanup** and **Next steps** (human-actionable only) share the 10 reserved lines.
@@ -118,6 +120,20 @@ plan file, and plan paths inside committed docs, are forbidden.
118
120
  and the docs cap-validator is green. An aborted plan still runs Cleanup — partial outputs land in
119
121
  `known_issues.md`.
120
122
 
123
+ ## The queue
124
+
125
+ `docs/plans/queue.md` NAMES work; it never holds the analysis of it. A row is one plain sentence
126
+ saying what the work is and for whom, then its id, then a short body — measurements, `file:line`
127
+ citations and fix direction belong to an ADR or the record the row points at. A row that goes
128
+ terminal is DELETED in the same change that mints its closing artifact — and where the queue is
129
+ gitignored, the deleted text is first written to a purge archive beside the history docs, because
130
+ there git history is no tombstone. In a hidden deployment that archive is machine-local like the rest
131
+ of the substrate: its guarantee is the working copy, not git. Frozen work with a stated resume
132
+ condition is not terminal and stays, in its own bucket. Order inside a bucket IS priority.
133
+ A prose promise to trim later has been measured failing: a checker over the file is the rung, and it
134
+ is runnable — `node <kit>/tools/queue-audit-cli.mjs --check docs/plans/queue.md --section '<bucket
135
+ heading>' --max-rows <n> --max-row-lines <n>`, declared as a project gate once the queue has migrated.
136
+
121
137
  ## The plan must read cold
122
138
 
123
139
  The executing session sees the plan file and the repository, never the authoring conversation.
@@ -38,7 +38,7 @@ Slots: review
38
38
  a `create` row and a revision of a governed contract a `modify` row, both written here so they
39
39
  exist AT review.
40
40
  3. **Self-review** — apply *What gets cut*; fold by code (read and cite the `file:line`); update
41
- `queue.md` for a series.
41
+ `queue.md` for a series, to the shape *The queue* fixes.
42
42
  4. **review {recipe}** — Solo (self-review only) / Reviewed (one backend) / Council (both; you
43
43
  synthesize), as the resolved `review` recipe selects.
44
44
  5. **Fold + loop** — fold every finding and re-review; CLEAN is **0 blockers + 0 majors** from every
@@ -117,9 +117,11 @@ refuse case per rule and an accept case per kind; a refuse fixture yields exactl
117
117
 
118
118
  - **One entity.** A spec is the durable contract; the ephemeral plan is the delta vehicle; the spec
119
119
  revision lands with the code. There is no change-spec entity.
120
- - **Governing specs are plural.** A plan cites ZERO (nothing spec-covered touched legal during
121
- adoption), ONE or MANY governing specs one per touched spec-covered slice; a shared-module change
122
- cites the specs of every slice whose contract it can alter.
120
+ - **Governing specs are plural.** A plan cites ZERO naming the adoption state it relies on:
121
+ `not adopted` (no store) or `adopting` (no live contract yet), either with a recorded decline, or
122
+ `nothing spec-covered touched`; a bare zero is never a licence — ONE or MANY governing specs —
123
+ one per touched spec-covered slice; a shared-module change cites the specs of every slice whose
124
+ contract it can alter.
123
125
  - **Out of scope composes PER GOVERNING SLICE — there is no global union.** Each cited spec's
124
126
  exclusions bound only the work inside that slice, and the plan's non-goals restate them per slice.
125
127
  A cross-spec conflict is resolved by a spec REVISION BEFORE plan approval — never by silent