@afokapu/atdd-bun 0.9.3 → 0.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -55,7 +55,7 @@ registerEnforcementTest({ root: import.meta.dir + "/..", profiles: ["traceabilit
55
55
  | `topology` | feature decomposition and the plan, source, test and E2E locations. Missing source, tests and E2E suites are reported only when a run also selects `coder` or `tester`, so a repository in the PLAN stage can enforce `planner` and `topology` before RED |
56
56
  | `planner` | schemas for every plan artifact, graph integrity, the scoped planner rules |
57
57
  | `telemetry` | the telemetry tracking plan: item shape, path-mirrored identity and versioning under `telemetry/`, wagon ownership of logical artifacts, the per-acceptance telemetry decision, metric label cardinality, source `Telemetry:` references, raw-string and forbidden-property emission, the vendor-SDK boundary around the TelemetryPort, and telemetry tests that bind the acceptance and item, assert the exact identity on a captured sink, cover every required item, and exercise declared timing semantics |
58
- | `delivery` | the review record of each tranche under `docs/delivery/tranches/`: allowed author and reviewer models with recorded fallbacks, reviewer independence, every finding fixed, withdrawn after one dispute or ruled on by a human, every configured stage approved, and, at the gate, no change without a record and a merged head that contains exactly the approved commit. Inert until adopted |
58
+ | `delivery` | the review record of each tranche under `docs/delivery/tranches/`: allowed writer and reviewer models with recorded fallbacks, reviewer independence, every finding fixed, withdrawn after one dispute or ruled on by a human, every configured stage approved, and, at the gate, no change without a record and a merged head that contains exactly the approved commit. Inert until adopted |
59
59
  | `docs` | the documentation capability, including the generated journey view |
60
60
  | `coder`, `tester`, `security`, `architecture`, `metrics`, `runtime` | Bun source and test conventions |
61
61
  | `interlocking` | train/interlocking binding, infrastructure and route coverage |
@@ -111,11 +111,20 @@ release: { enabled: false }
111
111
 
112
112
  ### Delivery
113
113
 
114
- For programs delivered as tranches by a coordinator and persistent drivers, with headless authors
114
+ For programs delivered as tranches by a coordinator and persistent drivers, with headless writers
115
115
  and independent reviewers. Adopt it by naming `delivery` in `profiles:` (or, with no list, by adding
116
- a `delivery:` block); `agent init` then installs the delivery skill and its review contract. The
117
- adopting pull request is itself governed: it changes files outside the delivery root, so it carries
118
- its own tranche record, reviewed and `ready` like any other.
116
+ a `delivery:` block); `agent init` then installs the delivery skill. The adopting pull request is
117
+ itself governed: the skill and workflow it brings are outside the delivery root, so it carries its
118
+ own tranche record, reviewed and `ready` like any other. A later change to `atdd-bun.yaml` alone
119
+ needs no record: the integrity check reports any loosening for a human to approve.
120
+
121
+ The policy names, for each lifecycle stage (plan, red, green, refactor, final), who may write it and
122
+ who reviews it; either may be absent. By default two stages are reviewed: the plan, and the whole
123
+ change at its head (final). Red, green and refactor are written and held by their gates. How the
124
+ coordinator and drivers carry this out is the `delivery.operating-model` convention; what a reviewer
125
+ checks is `delivery.review`; how agents talk, through a local ntfy board with one topic per program,
126
+ tranche and review conversation (`atdd-bun chat`), is `delivery.board`, used only when `delivery.board` is
127
+ set. The skill only points to them.
119
128
 
120
129
  The records live with the program's reasoning, in the docs profile's `docs/delivery/` area:
121
130
 
@@ -140,31 +149,40 @@ Upgrading from 0.9.0: an `atdd-bun.yaml` field with the wrong type (a quoted num
140
149
  non-string list item, `.inf`) is now reported, and on the base branch it blocks every pull request,
141
150
  since the policy cannot be compared. Correct such fields on the base branch before upgrading.
142
151
 
152
+ Upgrading to 0.10: the policy names writers and reviewers per stage. The 0.9 stage names
153
+ (`plan_review`, `test_review`, `code_review`, `final_review` with `authors` and `reviewers`) are still
154
+ read, as `plan`, `red`, `refactor` and `final`, and 0.9 records that name review authors still count as
155
+ recorded work. Independence now defaults to `different-model`: a repository that relied on the old
156
+ `fresh-process` default sets it explicitly. A repository with no `stages:` gets the new default
157
+ operating model, without the integrity check reporting the change.
158
+
143
159
  Every key is optional; these are the defaults:
144
160
 
145
161
  ```yaml
146
162
  delivery:
147
163
  root: docs/delivery/tranches # one <tranche>/evidence.yaml per tranche, reports beside it
148
- require_record: true # at the gate, a change outside the root needs a tranche record
149
- multiplexer: herdr # the terminal multiplexer agents run in; any command name
150
- independence: fresh-process # or different-model; overridable per stage
164
+ require_record: true # at the gate, a change outside the root (other than atdd-bun.yaml alone) needs a tranche record
165
+ multiplexer: herdr # the terminal multiplexer humans watch agents in; any command name
166
+ # board: { url: http://127.0.0.1:2586 } # opt-in: agents talk through a local board; absent, there is none
167
+ independence: different-model # a reviewer's model wrote none of the work it reviews; or fresh-process
151
168
  stages: # models in preference order: the first, then recorded fallbacks
152
- plan_review: { authors: [codex], reviewers: [glm, claude] }
153
- test_review: { authors: [glm, claude], reviewers: [codex, claude] }
154
- code_review: { authors: [glm, claude], reviewers: [glm, claude] }
155
- final_review: { authors: [codex], reviewers: [codex, claude] }
169
+ plan: { writer: [codex, claude-opus], reviewer: [glm, claude-opus, codex] }
170
+ red: { writer: [glm, claude-sonnet, claude-opus, codex] }
171
+ green: { writer: [glm, claude-sonnet, claude-opus, codex] }
172
+ refactor: { writer: [glm, claude-sonnet, claude-opus, codex] }
173
+ final: { reviewer: [codex, glm, claude-opus] }
156
174
  fallback: { after_failures: 3, within_minutes: 10, when_exhausted: block } # or wait
157
- commands: {} # per model: { author: "...", review: "..." } overriding the skill's defaults
175
+ commands: {} # per model: { author: "...", review: "..." } overriding delivery.operating-model's defaults
158
176
  ```
159
177
 
160
178
  The profile checks the record, never the running agents. The generated CI sets
161
179
  `ATDD_DELIVERY_GATE`: `merge` on pull requests and the merge queue, where every record the branch
162
180
  changes must be `ready`, the branch may differ from its approved SHA only under the delivery root,
163
- and a change outside the root needs a record (`require_record`, default true); `post-merge` on a
181
+ and a change outside the root needs a record unless it touches only `atdd-bun.yaml` (`require_record`, default true); `post-merge` on a
164
182
  push to a protected branch (the generated workflow's push branches follow `protected_branches`), where the pushed commit must contain every approved SHA it brings in. Merge tranches with a
165
183
  merge commit: a squash or rebase merge writes a commit no reviewer saw, and the post-merge check
166
184
  fails on it. Moving the root, dropping a stage, relaxing a stage from `different-model` to
167
- `fresh-process`, adding an author or reviewer, moving a fallback model earlier in a list, turning off
185
+ `fresh-process`, no longer reviewing a stage, adding a writer or reviewer, moving a fallback model earlier in a list, turning off
168
186
  `require_record`, making fallback easier, or adding or changing a model's `commands` loosens the policy and is reported by the integrity check.
169
187
 
170
188
  The record's model and run identifiers are the driver's claims. The profile checks that they are
@@ -192,7 +210,7 @@ delivery skill.
192
210
 
193
211
  - the installed package differs from its published hashes;
194
212
  - the dependency is not an npm registry version;
195
- - a generated file (workflow, skills, the delivery review contract, instruction block, integrity test) was edited;
213
+ - a generated file (workflow, skills, instruction block, integrity test) was edited;
196
214
  - `atdd-bun.yaml` is looser than on the base branch (after the first explicit `profiles:` list,
197
215
  dropping a profile or the list counts).
198
216
 
@@ -4,7 +4,7 @@ kind: rule
4
4
  status: active
5
5
  name: The approved SHA is a commit in this repository
6
6
  statement: >-
7
- A ready record's approved_sha resolves to a commit in the repository's history, and every configured stage's last approval names a commit in approved_sha's history, each stage's containing the one before it in lifecycle order (REQUIRED).
7
+ A ready record's approved_sha resolves to a commit in the repository's history, every work entry names a commit in approved_sha's history, and every reviewed stage's last approval does too, each containing the one before it in lifecycle order (REQUIRED).
8
8
  terms:
9
9
  - term_id: approved_sha
10
10
  text: >-
@@ -0,0 +1,60 @@
1
+ schema_version: 1.1.0
2
+ rule_id: delivery.board
3
+ kind: policy
4
+ status: active
5
+ name: Agents talk through a local board, one topic per conversation
6
+ statement: >-
7
+ Where atdd-bun.yaml sets delivery.board, the coordinator, drivers, writers and reviewers talk through a local message board with atdd-bun chat, never by typing into another agent's pane: one topic for the program, one per tranche, one per review conversation, each agent launched with its identity and the only topics it may use, every message naming its sender and recipients. The board makes the work visible; the evidence record stays the only thing the delivery rules judge.
8
+ terms:
9
+ - term_id: topics
10
+ text: >-
11
+ Three levels, each its own topic, named by atdd-bun chat topic so every agent spells them alike. atdd-<program>: coordinator, drivers and humans; status only (started, blocked, ready, merged, a decision a human must take). atdd-<program>-<tranche>: the tranche's driver and writers; tasks and results. atdd-<program>-<tranche>-<stage>-<round>: the driver and one reviewer; the review request and the verdict. A rebuttal opens the next round's topic.
12
+ values:
13
+ program: atdd-bun chat topic <program>
14
+ tranche: atdd-bun chat topic <program> <tranche>
15
+ review: atdd-bun chat topic <program> <tranche> <stage> <round>
16
+ - term_id: identity
17
+ text: >-
18
+ <role>@<tranche> for the whole tranche (driver@auth, writer-glm@auth, reviewer-codex@auth; a later round adds it: reviewer-codex-2@auth); coordinator for the coordinator.
19
+ - term_id: launch
20
+ text: >-
21
+ Whoever starts an agent gives it, in its environment, ATDD_AGENT (its identity) and ATDD_TOPICS (the topics it may use, comma-separated), and ATDD_BOARD_URL when the board is not the configured one. The coordinator gives a driver its tranche topic and the program topic; the driver adds each review topic as it opens it, gives a writer the tranche topic, and gives a reviewer its review topic alone. atdd-bun chat refuses any other topic, to read or to write.
22
+ values:
23
+ driver: ATDD_AGENT=driver@<tranche> ATDD_TOPICS=atdd-<program>,atdd-<program>-<tranche>
24
+ writer: ATDD_AGENT=writer-<model>@<tranche> ATDD_TOPICS=atdd-<program>-<tranche>
25
+ reviewer: ATDD_AGENT=reviewer-<model>@<tranche> ATDD_TOPICS=atdd-<program>-<tranche>-<stage>-<round>
26
+ - term_id: message
27
+ text: >-
28
+ atdd-bun chat post writes the header: from, to, participants and conversation always; kind (task, result, review-request, verdict, rebuttal, status, blocked), stage, reply-to, repo, branch, worktree, goal, base and sha when they apply. The body carries the whole instruction or answer, however long: the thread is the history.
29
+ - term_id: commands
30
+ text: >-
31
+ Post with the body on standard input; read what is on a topic; wait for the next message addressed to you, bounded by --timeout so a tool call ends (exit 2: nothing yet, wait again with --since the last id); and, for a human, atdd-bun chat alone opens the board: every topic on the left, the selected conversation on the right, live.
32
+ values:
33
+ post: atdd-bun chat post <topic> --to <agent> --kind <kind> [--reply-to <id>] < message.md
34
+ read: atdd-bun chat read <topic> [--mine]
35
+ wait: atdd-bun chat wait <topic> [--since <id>] --timeout 540
36
+ board: atdd-bun chat
37
+ - term_id: server
38
+ text: >-
39
+ A local ntfy server, bound to the loopback address, keeping its messages in a cache file. Every agent on the machine reaches it from the same address, so ntfy's per-client limits are raised: requests, subscriptions, daily messages, and new topics (one topic per conversation creates many). With the defaults, ten parallel tranches exhaust them and posts are refused.
40
+ values:
41
+ docker: docker run -d --name atdd-board -p 127.0.0.1:2586:80 -v ~/.atdd-board:/var/cache/ntfy -e NTFY_CACHE_FILE=/var/cache/ntfy/cache.db -e NTFY_CACHE_DURATION=720h -e NTFY_MESSAGE_SIZE_LIMIT=64k -e NTFY_VISITOR_REQUEST_LIMIT_BURST=100000 -e NTFY_VISITOR_REQUEST_LIMIT_REPLENISH=1ms -e NTFY_VISITOR_SUBSCRIPTION_LIMIT=10000 -e NTFY_VISITOR_MESSAGE_DAILY_LIMIT=0 -e NTFY_VISITOR_TOPIC_CREATION_LIMIT_BURST=0 binwiederhier/ntfy serve
42
+ content:
43
+ summary: >-
44
+ Typing into an agent's pane loses messages to startup screens, busy prompts and cut pastes. A board with one topic per conversation delivers every message, keeps the history, and lets a human follow any thread.
45
+ normative_text: |
46
+ A reviewer never reads the tranche's conversation: its brief is the review request on its own topic, and it posts its verdict there. The driver records that verdict in evidence.yaml; the board is never evidence.
47
+ Agents use atdd-bun chat only. It refuses an address that is not local and a topic the agent was not given. Never call ntfy directly: in the ntfy client -u means credentials, and a topic without a server goes to the public ntfy.sh.
48
+ A running session that must be woken is reached through its own channel (codex queue --thread <id>; --resume <session> for Claude and Pi), never by typing into its pane; its launch prompt tells it to wait on its topic between steps.
49
+ fix_hint: |
50
+ Derive the topic with atdd-bun chat topic, launch the agent with ATDD_AGENT and ATDD_TOPICS, and post with --to. To follow the work, run atdd-bun chat in a terminal pane. ntfy cannot list its topics, so the first message on a topic also posts the topic's name to the directory topic atdd-topics, which the view reads.
51
+ exceptions:
52
+ - >-
53
+ The board is opt-in: without a delivery.board block there is none, and atdd-bun chat refuses to post or read. With the block and no url it is at http://127.0.0.1:2586; ATDD_BOARD_URL moves it on one machine but never switches it on. Switching it on or off changes how agents talk, not what a record must prove, so it is not a loosening.
54
+ metadata:
55
+ aliases:
56
+ - DELIVERY-BOARD-001
57
+ introduced_in: 0.10.0
58
+ # A policy, not a rule: how agents talk cannot be validated from the repository. atdd-bun chat enforces its guard rails.
59
+ implementation:
60
+ type: none
@@ -4,31 +4,37 @@ kind: rule
4
4
  status: active
5
5
  name: The delivery policy in atdd-bun.yaml is well-formed
6
6
  statement: >-
7
- The `delivery:` block of atdd-bun.yaml validates against delivery-config.schema.json: known keys only, each stage listing at least one reviewer, models named in lowercase, independence one of fresh-process or different-model, the root in canonical form (no leading, trailing or doubled slash, no dot segment) not overlapping the plan, source, test, e2e or telemetry root, and, inside docs/, exactly docs/delivery/tranches (REQUIRED).
7
+ The `delivery:` block of atdd-bun.yaml validates against delivery-config.schema.json: known keys only, each stage naming a writer list, a reviewer list or both, models named in lowercase, independence one of fresh-process or different-model, the root in canonical form (no leading, trailing or doubled slash, no dot segment) not overlapping the plan, source, test, e2e or telemetry root, and, inside docs/, exactly docs/delivery/tranches (REQUIRED).
8
8
  terms:
9
9
  - term_id: policy
10
10
  text: >-
11
- the `delivery:` block: root, independence, stages (per stage: authors, reviewers, optional independence), fallback and commands. Every key is optional; an omitted key takes the package default.
11
+ the `delivery:` block: root, independence, stages (plan, red, green, refactor, final; per stage: writer, reviewer, optional independence), fallback and commands. Every key is optional; an omitted key takes the package default.
12
12
  content:
13
13
  summary: >-
14
14
  The skill and the validators read the same policy. A misspelled key would be ignored by one and not the other, so a malformed block is a finding, and the rest of the run judges evidence against the package defaults.
15
15
  normative_text: |
16
- The policy decides which models may author and review each stage, in preference order, how independent a reviewer must be, and when a model counts as unavailable. It is read by the delivery skill when it dispatches work and by this profile when it judges the record, so both must see the same, valid document.
16
+ The policy decides which models may write and review each stage, in preference order, how independent a reviewer must be, and when a model counts as unavailable. It is read by the delivery skill when it dispatches work and by this profile when it judges the record, so both must see the same, valid document.
17
17
  fix_hint: |
18
18
  Correct the key the finding names. The defaults are:
19
19
 
20
20
  delivery:
21
21
  root: docs/delivery/tranches
22
- independence: fresh-process
22
+ independence: different-model
23
23
  stages:
24
- plan_review: { authors: [codex], reviewers: [glm, claude] }
25
- test_review: { authors: [glm, claude], reviewers: [codex, claude] }
26
- code_review: { authors: [glm, claude], reviewers: [glm, claude] }
27
- final_review: { authors: [codex], reviewers: [codex, claude] }
24
+ plan: { writer: [codex, claude-opus], reviewer: [glm, claude-opus, codex] }
25
+ red: { writer: [glm, claude-sonnet, claude-opus, codex] }
26
+ green: { writer: [glm, claude-sonnet, claude-opus, codex] }
27
+ refactor: { writer: [glm, claude-sonnet, claude-opus, codex] }
28
+ final: { reviewer: [codex, glm, claude-opus] }
29
+
30
+ independence defaults to different-model. A stage may name a writer, a reviewer or both. The 0.9
31
+ names plan_review, test_review, code_review and final_review are still read, as plan, red, refactor
32
+ and final, with authors as writers and reviewers as reviewers.
28
33
  exceptions:
29
34
  - >-
30
- Adoption is not exempt from the gate: the pull request that adopts delivery changes files outside the delivery root,
31
- so it carries its own ready tranche record like any other change.
35
+ Adoption is not exempt from the gate: the pull request that adopts delivery brings the generated skill and workflow,
36
+ outside the delivery root, so it carries its own ready tranche record like any other change. A later change to
37
+ atdd-bun.yaml alone needs no record (delivery.merge-gate).
32
38
  - >-
33
39
  Inert until adoption: with no `delivery:` block and `delivery` absent from `profiles:`, no delivery rule emits anything.
34
40
  metadata:
@@ -4,19 +4,19 @@ kind: rule
4
4
  status: active
5
5
  name: Every tranche folder holds a well-formed evidence record
6
6
  statement: >-
7
- Every folder under the delivery root holds an evidence.yaml that validates against delivery-evidence.schema.json, whose tranche matches the folder name, whose reviews each name an author and a configured stage, and whose reports are data files (json, jsonl, yaml, yml, txt, md, log) inside the tranche's own folder, one per review; a ready record names one for every review (delivery.stages-complete); and the records folder holds nothing else: a data file belongs only as a report a record in its tranche names, and every other file, including a loose file directly under the root, is a finding (REQUIRED).
7
+ Every folder under the delivery root holds an evidence.yaml that validates against delivery-evidence.schema.json, whose tranche matches the folder name, whose work entries each name a writer and a stage the policy lets someone write, whose reviews each name a stage the policy reviews, and whose reports are data files (json, jsonl, yaml, yml, txt, md, log) inside the tranche's own folder, one per review; a ready record names one for every review (delivery.stages-complete); and the records folder holds nothing else: a data file belongs only as a report a record in its tranche names, and every other file, including a loose file directly under the root, is a finding (REQUIRED).
8
8
  terms:
9
9
  - term_id: tranche
10
10
  text: >-
11
11
  one independently mergeable piece of a program, delivered on its own branch by one persistent driver through PLAN, RED, GREEN, SMOKE, REFACTOR and TRACE, with its reviews.
12
12
  - term_id: evidence
13
13
  text: >-
14
- the tranche's record of its reviews at `<root>/<tranche>/evidence.yaml`: status (open or ready), base and approved SHAs, and the chronological list of reviews, each with stage, reviewed SHA, author, reviewer, fallbacks, verdict, the checklist it went through and its findings.
14
+ the tranche's record at `<root>/<tranche>/evidence.yaml`: status (open or ready), base and approved SHAs, the work (who wrote each stage, at which SHA, with fallbacks) and the chronological list of reviews, each with stage, reviewed SHA, reviewer, fallbacks, verdict, the checklist it went through and its findings.
15
15
  content:
16
16
  summary: >-
17
17
  The record is what CI can check. A review that is not written down, or written down in a shape the validator cannot read, did not happen as far as the merge gate is concerned.
18
18
  normative_text: |
19
- The profile judges the record a tranche leaves, not the agents while they work. Each review is appended, never edited, so a repair loop shows as a request-changes entry followed by a fresh review of the same stage. A review must name its author: independence is judged against the author, and without one it cannot be judged.
19
+ The profile judges the record a tranche leaves, not the agents while they work. Each review is appended, never edited, so a repair loop shows as a request-changes entry followed by a fresh review of the same stage. A review answers for the work written since the previous reviewed stage; independence is judged against those writers, so a review with no recorded writer cannot be judged. A 0.9 record names the writer as the review's author.
20
20
  fix_hint: |
21
21
  Create or repair `<root>/<tranche>/evidence.yaml`; the finding names the schema path at fault. A request-changes review carries at least one finding; every review carries its `checked` list.
22
22
  exceptions:
@@ -4,7 +4,7 @@ kind: rule
4
4
  status: active
5
5
  name: A tranche merges only its approved SHA
6
6
  statement: >-
7
- At the gate (ATDD_DELIVERY_GATE, which the generated CI sets to merge on pull requests and the merge queue and to post-merge on pushes to the protected branches), every evidence record the change touches is ready and approves a commit the head contains; before the merge the head differs from that commit only by the record and the reports it names; no record is deleted, and no record or report already on the base branch is modified (records 0.8.0 kept under delivery/ included, after the root moves); under the delivery root only records and the reports a changed record names change; and, with require_record (the default), a change outside the delivery root comes with a tranche record (REQUIRED).
7
+ At the gate (ATDD_DELIVERY_GATE, which the generated CI sets to merge on pull requests and the merge queue and to post-merge on pushes to the protected branches), every evidence record the change touches is ready and approves a commit the head contains; before the merge the head differs from that commit only by the record and the reports it names; no record is deleted, and no record or report already on the base branch is modified (records 0.8.0 kept under delivery/ included, after the root moves); under the delivery root only records and the reports a changed record names change; and, with require_record (the default), a change outside the delivery root comes with a tranche record, unless atdd-bun.yaml is the only such file (REQUIRED).
8
8
  terms:
9
9
  - term_id: merge_gate
10
10
  text: >-
@@ -30,6 +30,10 @@ content:
30
30
  generated workflow runs from the repository root, and drift from an approved commit is judged repository-wide.
31
31
  - >-
32
32
  require_record false exempts changes with no record; turning it off is a loosening the integrity check reports.
33
+ - >-
34
+ A change to atdd-bun.yaml alone needs no record: which models may write and review is a human's decision, and
35
+ the integrity check already reports any loosening for a human to approve; a tightening needs no review. A policy
36
+ change that comes with any other file outside the delivery root is reviewed with it.
33
37
  metadata:
34
38
  aliases:
35
39
  - DELIVERY-MERGE-GATE-001
@@ -4,11 +4,11 @@ kind: rule
4
4
  status: active
5
5
  name: Authors and reviewers are allowed models, and every fallback says why
6
6
  statement: >-
7
- Each review's author and reviewer model appears in that stage's authors and reviewers lists; a model after the first is used only with a recorded fallback from each model before it, stating the kind of unavailability (outage, rate_limit, no_report, timeout), the failures observed (at least delivery.fallback.after_failures) within a window no longer than delivery.fallback.within_minutes, and the reason (REQUIRED).
7
+ Each work entry's writer appears in that stage's writer list and each review's reviewer in its reviewer list; a model after the first is used only with a recorded fallback from each model before it, stating the kind of unavailability (outage, rate_limit, no_report, timeout), the failures observed (at least delivery.fallback.after_failures) within a window no longer than delivery.fallback.within_minutes, and the reason (REQUIRED).
8
8
  terms:
9
9
  - term_id: fallback
10
10
  text: >-
11
- a recorded switch from a preferred model to the next one in the stage's list, with role (author or reviewer) and a reason: an outage, a rate limit, or no auditable report. A REQUEST CHANGES verdict is never a reason.
11
+ a recorded switch from a preferred model to the next one in the stage's list, with role (writer or reviewer) and a reason: an outage, a rate limit, or no auditable report. A REQUEST CHANGES verdict is never a reason.
12
12
  content:
13
13
  summary: >-
14
14
  The lists are the operator's choice of who may do the work. A silent switch to another model hides an outage and can hide a reviewer chosen for being lenient.
@@ -18,7 +18,7 @@ content:
18
18
  Record the fallback on the review:
19
19
 
20
20
  fallback:
21
- - { role: reviewer, from: glm, kind: rate_limit, failures: 3, window: { from: "2026-09-25T09:00:00Z", to: "2026-09-25T09:08:00Z" }, reason: "429 on 3 attempts in 10 minutes" }
21
+ - { role: reviewer, from: codex, kind: rate_limit, failures: 3, window: { from: "2026-09-25T09:00:00Z", to: "2026-09-25T09:08:00Z" }, reason: "429 on 3 attempts in 10 minutes" }
22
22
 
23
23
  or use a model the stage lists.
24
24
  exceptions:
@@ -0,0 +1,91 @@
1
+ schema_version: 1.1.0
2
+ rule_id: delivery.operating-model
3
+ kind: policy
4
+ status: active
5
+ name: How a program is delivered in tranches, and who does what
6
+ statement: >-
7
+ A coordinator splits the program into tranches and keeps every slot busy; one persistent driver per tranche runs its lifecycle stages in order (plan, red, green, refactor, final), dispatching each stage's writer and reviewer from the delivery policy in atdd-bun.yaml as fresh headless processes, and records who wrote each stage and every review in the tranche's evidence.yaml.
8
+ terms:
9
+ - term_id: coordinator
10
+ text: >-
11
+ Owns throughput, never the work. It splits the program into tranches with explicit dependencies and writes why, the scope and the split in docs/delivery/index.adoc; activates a tranche once its own dependencies have merged (one still waiting may run plan and its review, nothing after); starts one driver per tranche in its own worktree; keeps provider health for the program, telling every driver to skip a model that is down; intervenes when a tranche is idle without being blocked, or has deliverables and no commit after two hours; and hands BLOCKED disputed-finding to a human with both sides. It never writes, reviews or merges.
12
+ - term_id: driver
13
+ text: >-
14
+ Owns one tranche end to end. For each stage of delivery.stages, in order, it starts the stage's writer (the first model in its writer list, or the next after a recorded fallback) in the tranche worktree, commits red on its own before green starts, passes the stage's gate, and appends a work entry; for each stage with a reviewer it starts a review. It never reviews.
15
+ - term_id: review_run
16
+ text: >-
17
+ Every review is a fresh process in its own detached worktree at the exact SHA (git worktree add --detach), given the delivery.review convention, the stage, the base SHA, the SHA and the RED commit. Afterwards git status --porcelain must be empty and HEAD still the SHA, or the review is a REVIEWER_FAILURE and does not count. The raw output is kept in the tranche folder and named in the review's report.
18
+ - term_id: fallback
19
+ text: >-
20
+ After fallback.after_failures failures within fallback.within_minutes (outage, rate limit, no auditable report, timeout), the driver uses the next model in the list and records the fallback with its role, kind, failures and window. A request-changes verdict is never a failure. With the list exhausted, when_exhausted block emits BLOCKED provider-unavailable; wait keeps retrying the last model.
21
+ - term_id: repair
22
+ text: >-
23
+ For each finding of a request-changes review the writer either fixes it or writes one rebuttal with evidence; a fresh review of the same stage follows. If that reviewer upholds a disputed finding, the driver emits BLOCKED disputed-finding and never disputes it again. Any commit, regenerated file, conflict fix or rebase after an approval cancels it and the change goes back to that stage's review.
24
+ - term_id: ready
25
+ text: >-
26
+ When the last reviewed stage approves, the driver sets status ready and approved_sha to that SHA, commits the evidence and its reports alone, pushes, and merges with a merge commit once CI is green. A squash or rebase merge writes a commit no reviewer saw, and CI fails it after the merge.
27
+ - term_id: events
28
+ text: >-
29
+ One line each, for the coordinator to wait on: PROGRAM_EVENT <tranche> <STAGE <stage> <sha> | WORKER_START <role> <model> <sha> | WORKER_END <role> <model> <verdict> | FALLBACK <role> <from>→<to> <reason> | PR_OPENED <url> | MERGED <sha> | BLOCKED <reason> | HEARTBEAT>.
30
+ - term_id: commands
31
+ text: >-
32
+ The default headless command per model, overridden per model under delivery.commands.<model>.author (the writer's command) and .review. {prompt} and {worktree} are substituted. Claude runs take project settings only, so a user's own allowances cannot widen a reviewer. glm runs through Pi, which has no permission system: as a reviewer it gets the read tool alone (the driver puts the diff in its prompt and records its output), and as a writer it runs only inside a container or sandbox limited to the worktree.
33
+ values:
34
+ codex:
35
+ author: codex exec --cd {worktree} --sandbox workspace-write "{prompt}"
36
+ review: codex exec --cd {worktree} --sandbox read-only "{prompt}"
37
+ claude-opus:
38
+ author: cd {worktree} && claude --model opus --setting-sources project --permission-mode acceptEdits --allowedTools Read Grep Glob Edit Write "Bash(bun:*)" "Bash(git:*)" --output-format json -p "{prompt}"
39
+ review: cd {worktree} && claude --model opus --setting-sources project --allowedTools Read Grep Glob "Bash(git show:*)" "Bash(git diff:*)" "Bash(git log:*)" --disallowedTools Edit Write NotebookEdit --output-format json -p "{prompt}"
40
+ claude-sonnet:
41
+ author: cd {worktree} && claude --model sonnet --setting-sources project --permission-mode acceptEdits --allowedTools Read Grep Glob Edit Write "Bash(bun:*)" "Bash(git:*)" --output-format json -p "{prompt}"
42
+ review: cd {worktree} && claude --model sonnet --setting-sources project --allowedTools Read Grep Glob "Bash(git show:*)" "Bash(git diff:*)" "Bash(git log:*)" --disallowedTools Edit Write NotebookEdit --output-format json -p "{prompt}"
43
+ glm:
44
+ author: cd {worktree} && pi -p --no-session --no-extensions --no-skills --no-context-files --tools read,bash,edit,write "{prompt}"
45
+ review: cd {worktree} && pi -p --no-session --no-extensions --no-skills --no-context-files --tools read "{prompt}"
46
+ - term_id: example_record
47
+ text: >-
48
+ A tranche in progress: the plan written by codex and approved by glm; red, green and refactor written by glm; the final review requesting changes. Only the finding's outcome is ever added to an earlier entry.
49
+ values:
50
+ tranche: api
51
+ status: open
52
+ base_sha: 3f2a91c
53
+ work:
54
+ - { stage: plan, sha: 8b10e44, writer: { model: codex, run: codex-plan-1 } }
55
+ - { stage: red, sha: 91aa0b2, writer: { model: glm, run: glm-red-1 } }
56
+ - { stage: green, sha: a4c0f11, writer: { model: glm, run: glm-green-1 } }
57
+ - { stage: refactor, sha: a4c0f11, writer: { model: glm, run: glm-refactor-1 } }
58
+ reviews:
59
+ - stage: plan
60
+ sha: 8b10e44
61
+ reviewer: { model: glm, run: glm-plan-review-1 }
62
+ verdict: approve
63
+ checked: [ACC-API-001, planner.decomposition]
64
+ report: docs/delivery/tranches/api/plan-1.json
65
+ - stage: final
66
+ sha: a4c0f11
67
+ reviewer: { model: codex, run: codex-final-1 }
68
+ verdict: request_changes
69
+ checked: [ACC-API-001, src/wagons/api, coder.bun.error-response-*]
70
+ findings:
71
+ - { id: F1, severity: high, evidence: "src/wagons/api/handler.ts:42", invariant: "coded error bodies", affects: [ui], proposed_fix: "return { code: 'API_NOT_FOUND' }" }
72
+ report: docs/delivery/tranches/api/final-1.json
73
+ content:
74
+ summary: >-
75
+ The policy says who may write and review each stage; this convention says how the coordinator and drivers carry it out. The record they leave is what the delivery rules judge.
76
+ normative_text: |
77
+ With delivery.board set, tasks, results, review requests and verdicts travel on the board (delivery.board), each agent launched with its identity and topics. Without it, a task is the prompt an agent is launched with and its result is the run's output, which the driver keeps. Agents run in panes of the multiplexer named in delivery.multiplexer (default herdr) only for humans to watch; learn its commands from its own help before the first dispatch. Never deliver work to an agent by typing into its pane.
78
+ Before a hosted model receives private repository content, confirm the user or organization authorized it.
79
+ Never edit the delivery skill, these conventions, or loosen the delivery policy to get a tranche through. If the policy must change, stop and ask the human.
80
+ fix_hint: |
81
+ Follow the driver's order: writer, gate, work entry; review, repair, fresh review; ready and merge commit. The example_record term shows a record the profile accepts.
82
+ exceptions:
83
+ - >-
84
+ A stage the policy does not name has no writer and no review; a stage with a writer but no reviewer is held by its gate alone.
85
+ metadata:
86
+ aliases:
87
+ - DELIVERY-OPERATING-MODEL-001
88
+ introduced_in: 0.10.0
89
+ # A policy, not a rule: how agents carry out the delivery policy cannot be validated. The record they leave is.
90
+ implementation:
91
+ type: none
@@ -0,0 +1,42 @@
1
+ schema_version: 1.1.0
2
+ rule_id: delivery.review
3
+ kind: policy
4
+ status: active
5
+ name: What each review establishes, and how a reviewer works
6
+ statement: >-
7
+ A reviewer is a fresh, read-only process for one stage of one tranche, given the stage, the base SHA, the SHA and the RED commit. It reviews the change between the base and the SHA, working through its stage's checklist systematically, systemically and adversarially, and returns one review entry with verdict, checked and findings.
8
+ terms:
9
+ - term_id: plan
10
+ text: >-
11
+ Are we building the right thing? The plan files the change adds or modifies: the decomposition (wagon, WMBT, acceptance, train, journey, contract) covers the intent; every acceptance is testable with one observable outcome and every WMBT has a SMOKE acceptance; edge cases, migrations, security and architecture constraints are represented; the acceptances can prove the feature; owned files do not overlap other tranches.
12
+ - term_id: final
13
+ text: >-
14
+ Does the change satisfy the approved plan? The tests, implementation and refactoring the change brings are read together. Every acceptance the plan names has a test bound by URN that asserts the behaviour, not a mock; each acceptance test differs from its RED commit only by the removed RED marker and Phase header, and any other change is justified; behaviour is correct on edge and error paths; the changed code keeps layering, composition, DTO, error-response and security sound; SMOKE runs through the real entry point; nothing drifts from the plan.
15
+ - term_id: red
16
+ text: >-
17
+ Where the policy gives red a reviewer. Every acceptance has a RED test bound by URN that fails for the missing behaviour and would still fail for a wrong implementation, asserting observable output rather than mocks.
18
+ - term_id: green_refactor
19
+ text: >-
20
+ Where the policy gives green or refactor a reviewer. Every behaviour is correct against its acceptance on edge and error paths; the changed code keeps layering, composition, DTO, error-response and security sound; SMOKE runs through the real entry point.
21
+ - term_id: method
22
+ text: >-
23
+ Systematic, the whole checklist, with everything checked listed in checked, not only what failed. Systemic, from the change to what it affects (callers, contracts, other tranches, downstream owners), named in affects; code the change does not touch or affect is out of scope. Adversarial, assuming the change is wrong and proving it; a finding needs a file and line, a failing command or a counter-example. A small change gets a short review.
24
+ content:
25
+ summary: >-
26
+ The two default reviews carry the judgement no gate can: the plan before RED, and the whole change at its head. The gates between them are deterministic.
27
+ normative_text: |
28
+ The reviewer runs in a detached worktree at the SHA and reads the change with git diff base..SHA. The gates are CI's, not the reviewer's: it does not re-run them, and judges what they cannot. It may read, search and use git show, git diff and git log. It never edits, commits or writes, including output-file options such as git diff --output=; the driver checks the worktree afterwards, and a review that wrote does not count (delivery.reviewer-independent). It judges and proposes; the author applies.
29
+ A finding carrying the author's rebuttal is withdrawn, left out, when the rebuttal holds, or upheld, repeated with the same id; there is no second round (delivery.findings-resolved).
30
+ The reviewer returns exactly one YAML document, the review entry of delivery-evidence.schema.json without fallback or report, which the driver adds: stage, sha, verdict (approve only with no critical or high finding), checked, and findings, each with id, severity (critical, high, medium, low), evidence, invariant, affects and proposed_fix, a precise description or short snippet, never a rewrite.
31
+ fix_hint: |
32
+ Give the reviewer this convention, the stage, the base SHA, the SHA and the RED commit. A review that skipped part of its checklist shows it in checked; run the stage again with a fresh reviewer.
33
+ exceptions:
34
+ - >-
35
+ A stage the policy does not name has no review; its checklist applies only where the policy names it.
36
+ metadata:
37
+ aliases:
38
+ - DELIVERY-REVIEW-001
39
+ introduced_in: 0.10.0
40
+ # A policy, not a rule: what a reviewer judges cannot be validated. The record it leaves is, by the delivery rules.
41
+ implementation:
42
+ type: none
@@ -2,26 +2,26 @@ schema_version: 1.1.0
2
2
  rule_id: delivery.reviewer-independent
3
3
  kind: rule
4
4
  status: active
5
- name: Every reviewer is a fresh process independent of the authors
5
+ name: Every reviewer is a fresh process independent of the writers
6
6
  statement: >-
7
- A reviewer's run never authored anything in the tranche and never reviewed another entry; where the stage requires different-model independence, the reviewer's model also differs from the stage author's (REQUIRED).
7
+ A reviewer's run never wrote anything in the tranche and never reviewed another entry; where the stage requires different-model independence (the default), the reviewer's model wrote none of the work the review answers for: its own stage and every stage since the previous reviewed one (REQUIRED).
8
8
  terms:
9
9
  - term_id: run
10
10
  text: >-
11
11
  the identity of one process: a session id, a pane id with its start time, or the retained report. Two entries with the same run were produced by the same process.
12
12
  - term_id: independence
13
13
  text: >-
14
- fresh-process: the reviewer is a new process that authored nothing. different-model: additionally, its model is not the author's.
14
+ fresh-process: the reviewer is a new process that wrote nothing. different-model: additionally, its model wrote none of the work it reviews; the final review answers for red, green and refactor when none of them is reviewed.
15
15
  content:
16
16
  summary: >-
17
- A reviewer that edits becomes an author, and a reviewer that already saw an earlier round is anchored to it. Independence is what makes an approval evidence.
17
+ A reviewer that edits becomes a writer, and a reviewer that already saw an earlier round is anchored to it. Independence is what makes an approval evidence.
18
18
  normative_text: |
19
- Every review is a new process. The run identifiers make that checkable: a reviewer run that appears as an author anywhere in the tranche, or on a second review, is not independent. The stricter mode is chosen per stage in the policy.
19
+ Every review is a new process. The run identifiers make that checkable: a reviewer run that appears as a writer anywhere in the tranche, or on a second review, is not independent. The stricter mode is chosen per stage in the policy.
20
20
  fix_hint: |
21
- Re-run the review in a fresh process and record its own run; under different-model, use a reviewer model other than the author's.
21
+ Re-run the review in a fresh process and record its own run; under different-model, use the next reviewer model in the list that wrote none of the reviewed work.
22
22
  exceptions:
23
23
  - >-
24
- The persistent driver authors the plan and opens the PR; its run may never review.
24
+ The persistent driver dispatches the work and opens the PR; its run may never review.
25
25
  metadata:
26
26
  aliases:
27
27
  - DELIVERY-REVIEWER-INDEPENDENT-001
@@ -4,11 +4,11 @@ kind: rule
4
4
  status: active
5
5
  name: A ready record has every configured stage approved
6
6
  statement: >-
7
- A record with status ready has, for every stage the policy configures, a last review that approves; its approved_sha is the SHA the last review of the closing stage approved; every review names its retained raw report, which exists; and, when code_review is configured and is not the closing stage, nothing outside the delivery root changed between code_review's last approval and approved_sha, since a code change goes back through code_review (REQUIRED).
7
+ A record with status ready has, for every stage the policy gives a writer, recorded work, and for every stage it gives a reviewer, a last review that approves; its approved_sha is the SHA the last review of the closing stage approved; and every review names its retained raw report, which exists (REQUIRED).
8
8
  terms:
9
9
  - term_id: closing_stage
10
10
  text: >-
11
- the last configured stage in lifecycle order (plan_review, test_review, code_review, final_review); by default final_review, the review of the PR head.
11
+ the last reviewed stage in lifecycle order (plan, red, green, refactor, final); by default final, the review of the PR head, which answers for everything written since the plan review.
12
12
  content:
13
13
  summary: >-
14
14
  Ready is the driver's claim that the tranche can merge. It is only true when no stage is missing and the approval covers the commit being merged.
@@ -1,5 +1,5 @@
1
- # Ready, but final_review never ran, the approved SHA exists nowhere, a reviewer is off-list, and one
2
- # reviewer also authored.
1
+ # Ready, but the approved SHA is not what final_review approved and exists nowhere, a reviewer is off-list,
2
+ # one reviewer also authored, and a finding was never resolved.
3
3
  tranche: api
4
4
  status: ready
5
5
  base_sha: 3f2a91c
@@ -11,16 +11,10 @@ reviews:
11
11
  reviewer: { model: gpt, run: gpt-plan-1 }
12
12
  verdict: approve
13
13
  checked: [ACC-API-001]
14
- - stage: test_review
14
+ - stage: final_review
15
15
  sha: 91aa0b2
16
- author: { model: glm, run: glm-red-1 }
16
+ author: { model: codex, run: codex-green-1 }
17
17
  reviewer: { model: codex, run: driver-api }
18
- verdict: approve
19
- checked: [ACC-API-001]
20
- - stage: code_review
21
- sha: a4c0f11
22
- author: { model: glm, run: glm-green-1 }
23
- reviewer: { model: glm, run: glm-code-1 }
24
18
  verdict: request_changes
25
19
  checked: [ACC-API-001]
26
20
  findings:
@@ -29,9 +23,9 @@ reviews:
29
23
  evidence: src/wagons/api/handler.ts:42 returns a bare string
30
24
  invariant: every error response carries a coded body
31
25
  proposed_fix: return a coded error body
32
- - stage: code_review
33
- sha: "0000000"
34
- author: { model: glm, run: glm-green-2 }
35
- reviewer: { model: glm, run: glm-code-2 }
26
+ - stage: final_review
27
+ sha: a4c0f11
28
+ author: { model: codex, run: codex-green-2 }
29
+ reviewer: { model: codex, run: codex-final-2 }
36
30
  verdict: approve
37
31
  checked: [ACC-API-001, F1]