@sabaiway/agent-workflow-engine 3.0.0 → 3.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,45 @@ All notable changes to the methodology engine. Versions are this **package's** n
4
4
  they are distinct from the **deployment-lineage** stamp written into a project's `docs/ai/`
5
5
  (which tracks the shared `agent-workflow` lineage, head `3.0.0`).
6
6
 
7
+ ## 3.1.0 — a finding NAMES the invariant its fix enforces, and the acceptance criteria become a machine-readable list (AD-110)
8
+
9
+ A review round produces findings, and the canon never said which of them the phase owes. The two
10
+ ways of getting that wrong are silent and opposite: work the plan already requires gets deferred into
11
+ a queue row nobody reads, and work that was never in scope gets folded in until the round count is
12
+ the only thing that converges. `plan-execution` step 5 now carries the **finding-scope rule**.
13
+
14
+ - **`procedures.md` — the rule, in the plan-execution review step.** Every finding NAMES the
15
+ invariant its fix would enforce, BEFORE the edit, in every round, and where that invariant already
16
+ lives decides the arm. Already an acceptance criterion → **fold here**. It would have to be ADDED →
17
+ the **narrow fix** for the found site ships now (red first) and ONLY the generalization is queued,
18
+ as a row carrying five fields: the invariant, the origin `file:line`, the narrow fix, its proof,
19
+ and a residual exposure declared NOT live. No correct narrow fix → **blocking**: the phase does not
20
+ close, and it is never queued.
21
+ - **Two round bars ride along, declared before each round.** A finding counts only if it changes a
22
+ WRITE/REMOVE decision or is a false statement in shipped text; a repeat finding in one subarea
23
+ routes to SUBTRACTION, not a fourth patch.
24
+ - **Plan-execution scope only.** Plan-authoring settles boundaries and has no shipped behaviour to
25
+ call a live defect in, so it carries none of the three — pinned in both directions, so neither a
26
+ silent deletion nor a scope-creeping copy survives with tests green.
27
+ - **`planning.md` — the acceptance criteria ARE the `- ` bullets under `## Verification`, and they
28
+ are the whole list.** Nothing outside a bullet is one, and a claim matches WITHIN ONE bullet,
29
+ because bullets are reordered, split and deleted independently. A criterion needing two bullets is
30
+ two criteria. A Verification written as prose declares NO criteria and every finding against that
31
+ plan is a new invariant — the closed list fails closed.
32
+ - **The agent-rules lens gains the rule as one bullet, and that bullet carries its SCOPE.** The lens
33
+ intro applies every bullet to plan-AUTHORING as well, so an unqualified bullet would have
34
+ contradicted the canon it is rendered from — it opens `Finding scope (plan-execution)` and says
35
+ plan-review carries none of it. The outgoing body is appended to the append-only prior store, so
36
+ every deployed `agent_rules.md` still carrying the previous canonical body converges on first
37
+ touch.
38
+
39
+ Engine-only release: no migration, no structural change to a deployed `docs/ai/`, and the
40
+ deployment-lineage stamp does not move. One thing in a deployment DOES change, stamp-independently —
41
+ the `agent_rules.md` lens region itself, which the kit's `lens-region` reconcile refreshes on its
42
+ next touch for any deployment still carrying a known canonical body; a customized region is preserved
43
+ verbatim and flagged. The checker the rule names ships in **agent-workflow-kit 7.1.0**
44
+ (`fold-scope`).
45
+
7
46
  ## 3.0.0 — the plan canon becomes a capped index: a module ledger whose rows ARE the steps (AD-104)
8
47
 
9
48
  A plan is an **index plus constraints**, never a transcript. The executor reads the repository; the
package/SKILL.md CHANGED
@@ -3,7 +3,7 @@ name: agent-workflow-engine
3
3
  description: Canonical home of the agent-workflow planning methodology — the capped plan shape (goal and boundary, module ledger, verification), plan lifecycle, queue.md series index, mandatory Cleanup phase, the bounded methodology slot fragment, the orchestration-recipe vocabulary (Solo / Reviewed / Council / Delegated), and the activity-procedures canon (plan-authoring / plan-execution, with typed recipe slots). A published, installable npm package (available:true) that *provides* the methodology text; it mutates nothing. The composition root (agent-workflow-kit) reads this canon LIVE from the installed engine and injects the bounded slots from it — one source of truth, no bundled mirror; `npx @sabaiway/agent-workflow-kit@latest init` installs the engine.
4
4
  disable-model-invocation: true
5
5
  metadata:
6
- version: '3.0.0'
6
+ version: '3.1.0'
7
7
  ---
8
8
 
9
9
  # agent-workflow-engine
package/capability.json CHANGED
@@ -3,7 +3,7 @@
3
3
  "schema": 1,
4
4
  "name": "agent-workflow-engine",
5
5
  "kind": "methodology-engine",
6
- "version": "3.0.0",
6
+ "version": "3.1.0",
7
7
  "available": true,
8
8
  "provides": ["plan"],
9
9
  "roles": {},
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@sabaiway/agent-workflow-engine",
3
- "version": "3.0.0",
3
+ "version": "3.1.0",
4
4
  "description": "Canonical home of the agent-workflow planning methodology — the capped plan shape (goal and boundary, module ledger, verification), plan lifecycle, queue.md series index, and mandatory Cleanup phase, consumed by the kit (composition root). The methodology engine of the agent-workflow family.",
5
5
  "keywords": [
6
6
  "ai-agents",
@@ -143,3 +143,19 @@ Apply these when authoring a plan, reviewing, folding a finding, or editing code
143
143
  - **Recipe fidelity.** Council runs every backend the recipe names, **every round**; silently dropping a ready backend for quota/convenience is a forbidden downgrade — an unavailable backend is a LOUD, stated degrade, never a quiet drop.
144
144
  - **ExitPlanMode ≠ execute.** A harness "approved — start coding" prompt authorizes the PLAN only; this methodology overrides it. Continue into execution only as a DELIBERATE transition after the plan + cold-start prompt exist, never an implicit slide.
145
145
  - **Cost lanes.** Route every step to the **cheapest adequate executor** — L0 deterministic script (the batched gate matrix over `gates.json`, the rotation `--check`s) · L1 cheap subagent (extraction/drafting only; the orchestrator verifies) · L2 subscription bridge · L3 frontier judgment. A step with **no named guardrail does not move down** a lane, and the **red lines never move down** (council review models · real code · ADR/handover/changelog-entry wording · persuasive copy · go/no-go · the approval asks). Own-error repair: salvage recorded state first (L0/L1, batched), never frontier re-derivation. **Prompt economy:** read-only fan-out (research/sweeps/extraction) runs ONLY on restricted-tool vehicles — a full-tool subagent for read-only work is a forbidden lane downgrade (invisible prompt-flood + blast radius), and a subagent is never told to shell out for facts obtainable read-only; the orchestrator's own shell form is ONE plain pipeline per call (a `;`/`&&` chain or env-prefixed invocation never matches a prefix allow rule); a fan-out launcher that gates per call yields to the agent-spawn lane — capability-gated: without restricted-tool vehicles (generic full-tool spawning does not count), read-only research stays in the orchestrator's own context, never a vehicle mandate a host cannot satisfy. Judgment, code, synthesis stay at the frontier lane (a task that genuinely runs/writes keeps a full-tool subagent); honest limit: no deterministic gate classifies a dispatch — canon at the point of use + placed vehicles + the retro loop. **Writer economy:** a stage's repeated WRITER commands batch — evidence declarations ride consecutive plain invocations of ONE allow-listed tool, other stage writers combine via one launcher per stage; never an unbatched writer scatter (each gated write is its own prompt).
146
+
147
+ <!-- prior: 2026-08-22 (the finding-scope rule) — the outgoing body before the fold-channel bullet -->
148
+ ### 2.x. Planning, review & process-fidelity invariants
149
+ Apply these when authoring a plan, reviewing, folding a finding, or editing code — the layer read **before any code change**. (Full canon: the project's planning / workflow-methodology + orchestration canon. This section is rendered from that canon and refreshed on upgrade; a custom edit is preserved verbatim, but flagged.)
150
+ - **Fold by code, not prose.** Before folding a code-touching finding into a plan or change, read the cited `file:line` and cite it — a prose fold drifts from the code and seeds the next bug.
151
+ - **Right altitude.** Pin intent + invariants + acceptance criteria (named tests); leave fine code-mechanics to Execute, where prose cannot diverge from reality.
152
+ - **No code-mechanics in the plan.** A ledger row carries its path and anchor, and Verification carries the exact commands (the plan-shape canon) — checked syntax: the plan's own Verification runs them against an explicit expected outcome or gate; the only other syntax a plan may carry is a literal fixture/schema fragment a named test copies or validates. Un-run, logic-bearing syntax — control-flow, a regex, a glob, a grammar, an algorithm body, a mini-DSL — never lives in plan prose, however plausible or shell-verified it looks: a fold or draft that wants one is the trigger to write the test instead.
153
+ - **Test-as-spec.** Fold a code-touching finding into a red→green TEST, not a prose paragraph — the gate is the only deterministic checker; a paragraph cannot self-check.
154
+ - **Characterize-first.** Before editing UNCOVERED code, pin its current behavior in a green test, then edit — any unintended change goes red. Never edit what has no checker; first give it one. Keep edits atomic/reversible; prefer SUBTRACTIVE folds.
155
+ - **Fold minimally — prose has no checker.** An ephemeral, gitignored plan is prose with no executable checker; fold **minimally, in ONE place** and run a **self-consistency** read across the plan before every re-review — a fold that drifts several prose spots is what turns a 2-round review into churn.
156
+ - **Heavy review at the diff.** Plan-review settles architecture only (≤2 rounds, stop at the pre-existing→fold-induced crossover); the exhaustive per-row review runs against real compiling code + the full suite, where a regression fails a gate immediately. **Backend divergence** (one backend grounded-ships while another keeps revising mechanics) IS that crossover — resolve at altitude, don't exhaust the strictest backend; route an all-mechanics/CI or prose-only artifact to a **thin plan + diff-review**.
157
+ - **Convergence bar.** A review loop is CLEAN only when one round returns **0 blockers + 0 majors** from EVERY backend the recipe names (nits + a ship verdict is the stop). Folding ≠ convergence — re-review after folding.
158
+ - **Per-round emission.** Every review round emits **{round N · finding-origin tally · per-backend verdict}** so the crossover is a computed, visible signal, not a remembered rule.
159
+ - **Recipe fidelity.** Council runs every backend the recipe names, **every round**; silently dropping a ready backend for quota/convenience is a forbidden downgrade — an unavailable backend is a LOUD, stated degrade, never a quiet drop.
160
+ - **ExitPlanMode ≠ execute.** A harness "approved — start coding" prompt authorizes the PLAN only; this methodology overrides it. Continue into execution only as a DELIBERATE transition after the plan + cold-start prompt exist, never an implicit slide.
161
+ - **Cost lanes.** Route every step to the **cheapest adequate executor** — L0 deterministic script (the batched gate matrix over `gates.json`, the rotation `--check`s) · L1 cheap subagent (extraction/drafting only; the orchestrator verifies) · L2 subscription bridge · L3 frontier judgment. A step with **no named guardrail does not move down** a lane, and the **red lines never move down** (council review models · real code · ADR/handover/changelog-entry wording · persuasive copy · go/no-go · the approval asks). Own-error repair: salvage recorded state first (L0/L1, batched), never frontier re-derivation. **Prompt economy:** read-only fan-out (research/sweeps/extraction) runs ONLY on restricted-tool vehicles — a full-tool subagent for read-only work is a forbidden lane downgrade (invisible prompt-flood + blast radius), and a subagent is never told to shell out for facts obtainable read-only; the orchestrator's own shell form is ONE plain pipeline per call (a `;`/`&&` chain or env-prefixed invocation never matches a prefix allow rule); a fan-out launcher that gates per call yields to the agent-spawn lane — capability-gated: without restricted-tool vehicles (generic full-tool spawning does not count), read-only research stays in the orchestrator's own context, never a vehicle mandate a host cannot satisfy. Judgment, code, synthesis stay at the frontier lane (a task that genuinely runs/writes keeps a full-tool subagent); honest limit: no deterministic gate classifies a dispatch — canon at the point of use + placed vehicles + the retro loop. **Writer economy:** a stage's repeated WRITER commands batch — evidence declarations ride consecutive plain invocations of ONE allow-listed tool, other stage writers combine via one launcher per stage; never an unbatched writer scatter (each gated write is its own prompt).
@@ -1,6 +1,7 @@
1
1
  ### 2.x. Planning, review & process-fidelity invariants
2
2
  Apply these when authoring a plan, reviewing, folding a finding, or editing code — the layer read **before any code change**. (Full canon: the project's planning / workflow-methodology + orchestration canon. This section is rendered from that canon and refreshed on upgrade; a custom edit is preserved verbatim, but flagged.)
3
3
  - **Fold by code, not prose.** Before folding a code-touching finding into a plan or change, read the cited `file:line` and cite it — a prose fold drifts from the code and seeds the next bug.
4
+ - **Finding scope (plan-execution) — name the invariant BEFORE the edit.** During EXECUTION only — a plan under authoring has no shipped behaviour to call a live defect in, so plan-review carries none of this. Every finding names the invariant its fix would enforce, and where that invariant already lives decides the disposition: already an acceptance criterion of the phase → **fold here**; it would have to be ADDED → ship the **narrow fix** for the found site (red first, then green) and queue ONLY the generalization — a deferral row carries the invariant, the origin `file:line`, the narrow fix, its proof and a residual exposure declared NOT live; no correct narrow fix → **blocking**: the phase does not close, and it is **never queued**. Two bars declared before each round: a finding counts only if it changes a **WRITE/REMOVE decision** or is a false statement in shipped text; a repeat finding in one subarea **routes to SUBTRACTION**, not a fourth patch.
4
5
  - **Right altitude.** Pin intent + invariants + acceptance criteria (named tests); leave fine code-mechanics to Execute, where prose cannot diverge from reality.
5
6
  - **No code-mechanics in the plan.** A ledger row carries its path and anchor, and Verification carries the exact commands (the plan-shape canon) — checked syntax: the plan's own Verification runs them against an explicit expected outcome or gate; the only other syntax a plan may carry is a literal fixture/schema fragment a named test copies or validates. Un-run, logic-bearing syntax — control-flow, a regex, a glob, a grammar, an algorithm body, a mini-DSL — never lives in plan prose, however plausible or shell-verified it looks: a fold or draft that wants one is the trigger to write the test instead.
6
7
  - **Test-as-spec.** Fold a code-touching finding into a red→green TEST, not a prose paragraph — the gate is the only deterministic checker; a paragraph cannot self-check.
@@ -67,6 +67,14 @@ command over its `check-id`s — existence and budget for create/modify, absence
67
67
  for a sweep, and the total line. Per-row assertions in prose are the repetition this section exists
68
68
  to avoid.
69
69
 
70
+ **The acceptance criteria ARE the `- ` bullets.** Every top-level `- ` bullet in this section is one
71
+ acceptance criterion, and they are the whole list — nothing outside a bullet is one. That makes the
72
+ list machine-readable, so a review can be told mechanically whether a claimed invariant is already
73
+ required. A claim matches WITHIN ONE bullet: a literal spanning two is not in scope, because bullets
74
+ are reordered, split and deleted independently. A criterion that needs two bullets is two criteria —
75
+ write each one self-contained. A Verification written as prose with no bullets therefore declares NO
76
+ criteria, and every finding against that plan is a new invariant: the closed list fails closed.
77
+
70
78
  ## What gets cut
71
79
 
72
80
  Delete any line for which both answers are yes: *can a zero-context executor still pick the right
@@ -79,6 +79,15 @@ Each ledger row is one logical commit.
79
79
  `core-evidence red-proof` declares each bugfix red BEFORE the fix; `core-evidence
80
80
  degrade` records an unavailable backend; reviews run on the STAGED tree; `run-gates --final`
81
81
  mints the ONE receipt `commit-guard --check` gates the commit against.
82
+
83
+ **Finding scope** — every finding NAMES the invariant its fix enforces, BEFORE the edit, every
84
+ round. Already an acceptance criterion (*Verification*'s `- ` bullets) → **fold here**. It would
85
+ have to be ADDED → the **narrow fix** for the found site ships now (red first) and ONLY the
86
+ generalization is queued, as a row carrying five fields: the invariant, the origin `file:line`,
87
+ the narrow fix, its proof, and a residual exposure declared NOT live. No correct narrow fix →
88
+ **blocking** — the phase does not close, and it is never queued. Two bars, before each round: a
89
+ finding counts only if it changes a WRITE/REMOVE decision or is a false statement in shipped
90
+ text; a repeat finding in one subarea routes to SUBTRACTION, not a fourth patch.
82
91
  6. **Gates** — the project's verification gate to green.
83
92
  7. **Commit boundary** — the orchestrator makes the single commit; a backend never commits; the
84
93
  commit-approval policy lives in the project's own rules.