@afokapu/atdd-bun 0.7.2 → 0.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (34) hide show
  1. package/README.md +45 -4
  2. package/conventions/delivery/delivery.approved-sha-resolves.convention.yaml +33 -0
  3. package/conventions/delivery/delivery.config-schema.convention.yaml +43 -0
  4. package/conventions/delivery/delivery.evidence-schema.convention.yaml +34 -0
  5. package/conventions/delivery/delivery.findings-resolved.convention.yaml +42 -0
  6. package/conventions/delivery/delivery.merge-gate.convention.yaml +42 -0
  7. package/conventions/delivery/delivery.model-allowed.convention.yaml +36 -0
  8. package/conventions/delivery/delivery.reviewer-independent.convention.yaml +34 -0
  9. package/conventions/delivery/delivery.stages-complete.convention.yaml +31 -0
  10. package/detectors/delivery_evidence/atdd.implementation.yaml +33 -0
  11. package/detectors/delivery_evidence/detect.mjs +9 -0
  12. package/detectors/delivery_evidence/fixtures/clean/atdd-bun.yaml +8 -0
  13. package/detectors/delivery_evidence/fixtures/clean/delivery/api/evidence.yaml +54 -0
  14. package/detectors/delivery_evidence/fixtures/dirty/atdd-bun.yaml +3 -0
  15. package/detectors/delivery_evidence/fixtures/dirty/delivery/api/evidence.yaml +37 -0
  16. package/detectors/delivery_evidence/fixtures/dirty/delivery/data/evidence.yaml +8 -0
  17. package/detectors/delivery_evidence/fixtures/dirty/delivery/orphan/README.md +1 -0
  18. package/detectors/delivery_evidence/fixtures/dirty/delivery/ui/evidence.yaml +5 -0
  19. package/integrity.json +33 -11
  20. package/package.json +1 -1
  21. package/planner-schemas/delivery-config.schema.json +80 -0
  22. package/planner-schemas/delivery-evidence.schema.json +84 -0
  23. package/relationships.yaml +98 -0
  24. package/src/agent.ts +16 -2
  25. package/src/ci.ts +8 -1
  26. package/src/delivery.ts +415 -0
  27. package/src/enforce.ts +2 -1
  28. package/src/index.ts +2 -0
  29. package/src/integrity.ts +27 -10
  30. package/src/setup.ts +5 -2
  31. package/templates/agents/AGENTS.block.md +4 -2
  32. package/templates/agents/delivery/SKILL.md +77 -0
  33. package/templates/agents/delivery/review.md +45 -0
  34. package/templates/github/atdd-bun.yml +6 -1
@@ -0,0 +1,77 @@
1
+ ---
2
+ name: delivery
3
+ description: Use when a program is delivered as tranches by a coordinator and persistent drivers, with headless authors and independent reviewers. Covers activation, dispatch, model fallback, the four reviews, disputes and the evidence record the `delivery` profile checks in CI.
4
+ ---
5
+ <!-- Generated by @afokapu/atdd-bun {{VERSION}}. Do not edit; run `atdd-bun agent init --replace`. -->
6
+
7
+ The policy is the `delivery:` block of `atdd-bun.yaml`: per stage, the allowed authors and reviewers in preference order, the independence mode, and the fallback thresholds. Read it; do not hard-code models. The rules are `conventions/delivery/*.convention.yaml` in `node_modules/@afokapu/atdd-bun/`. Gate: `bun run atdd-bun delivery`.
8
+
9
+ A **tranche** is one independently mergeable piece of the program, on its own branch and worktree, owned end to end by one persistent **driver**. Each tranche runs the ATDD lifecycle (`.agents/skills/atdd/SKILL.md`) with four reviews: `plan_review` after PLAN, `test_review` after RED, `code_review` after GREEN → SMOKE → REFACTOR → TRACE, and `final_review` of the PR head.
10
+
11
+ ## Coordinator
12
+
13
+ Its job is throughput: every worker slot busy, every tranche moving. It never implements, repairs tests, reviews, or merges a tranche.
14
+
15
+ 1. Split the program into tranches with explicit dependencies. Activate a tranche as soon as its own dependencies have merged; do not wait for a whole wave. A tranche whose dependencies are still open may run PLAN and `plan_review` but nothing after; revalidate its plan once they merge.
16
+ 2. For each active tranche, create a worktree from the owning repository's workspace and start a driver in it through the multiplexer (see Multiplexer). Send the mandate, submit it, wait 4–6 s, and read the pane: a working indicator or agent output means it landed; an empty prompt or placeholder means retry before waiting on anything.
17
+ 3. Wait on the multiplexer's events, not polling loops, and on every driver at once. When a slot frees, give it to the next ready tranche or to planning ahead.
18
+ 4. Keep provider health for the whole program. When a driver reports a model unavailable, tell every driver to go straight to the next model in its lists until it recovers, so no tranche spends time rediscovering an outage.
19
+ 5. Intervene when a tranche is not BLOCKED yet no worker has run for a while, or when two hours pass with deliverable-shaped changes and no commit: tell the driver to commit, push and open its PR.
20
+ 6. A `BLOCKED disputed-finding` goes to the human with both sides. `BLOCKED provider-unavailable` frees the slot for other work.
21
+
22
+ ## Multiplexer
23
+
24
+ Agents run in panes of the terminal multiplexer named in `delivery.multiplexer` (default `herdr`); say which one you are using when you start. Do not assume its commands: before the first dispatch, read its own help (`<multiplexer> --help`, then `<multiplexer> <group> --help`, and any schema it publishes) and map each operation below to a command. If one is missing, report it instead of scripting around it.
25
+
26
+ | Operation | herdr |
27
+ |---|---|
28
+ | create a worktree under the owning repository's workspace | `herdr worktree create --workspace <repo-ws> --branch <b> --base <sha> --path <p> --label <tranche> --no-focus` |
29
+ | start an agent in a pane with a working directory | `herdr agent start <name> --cwd <worktree> --workspace <ws> --no-focus -- codex` |
30
+ | send text, then submit it | `herdr agent send <name> "<text>"`, then `herdr pane send-keys <pane> Enter` (send does not press Enter) |
31
+ | read recent output | `herdr agent read <pane> --source recent-unwrapped --lines 12` |
32
+ | wait for one pane's status or output | `herdr wait agent-status <pane> --status idle`, `herdr wait output <pane> --match "PROGRAM_EVENT" --regex` |
33
+ | wait on every pane at once | the socket API's `events.subscribe` (`herdr api schema --json`): `pane_agent_status_changed`, `pane_exited`, `pane_output_changed` |
34
+ | inspect a pane's process | `herdr pane process-info --pane <pane>` |
35
+
36
+ ## Driver
37
+
38
+ 1. Run each stage's author and each review as a separate headless process, a review in its own pane so its output is retained. Build the command from `delivery.commands` or the defaults below. An author runs in the tranche worktree.
39
+ 2. Run every review in its own detached worktree at the exact SHA (`git worktree add --detach <path> <sha>`), never in the tranche worktree, and give it `.agents/skills/delivery/review.md`, the stage and the SHA. The reviewer only reads and proposes. Tool allowlists cannot make a CLI fully read-only (`git diff --output=` writes), so isolation does: afterwards `git -C <path> status --porcelain` must be empty and HEAD still the SHA, or the review is a `REVIEWER_FAILURE` and does not count. Remove the worktree after retaining the report.
40
+ 3. Fallback: after `fallback.after_failures` failures within `fallback.within_minutes` (outage, rate limit, no auditable report), use the next model in the stage's list and record it with its `kind` (`outage`, `rate_limit`, `no_report`, `timeout`), `failures`, and the `window` from the first to the last counted failure. REQUEST CHANGES is never a failure. With the list exhausted, `when_exhausted: block` emits `BLOCKED provider-unavailable`; `wait` keeps retrying the last model.
41
+ 4. For each finding of a REQUEST CHANGES review, either have the author fix it, or write one rebuttal with evidence (a test result, a rule id, file:line). Then run a fresh review of the same stage. If that reviewer upholds a disputed finding, emit `BLOCKED disputed-finding`; never dispute it a second time.
42
+ 5. Any commit, regenerated file, conflict fix or rebase after an approval cancels it. A change to the code goes back through `code_review`, then `final_review`.
43
+ 6. Append every review to `<root>/<tranche>/evidence.yaml` as it happens; never edit an earlier entry, only add each finding's `outcome` (`fixed`, `withdrawn`, or `human` with the `decision`). Keep the raw reviewer output inside the tranche's folder and name it in `report`; nothing else goes in that folder. When `final_review` approves, set `status: ready` and `approved_sha` to that SHA, commit the evidence and reports alone, push, and merge with a merge commit once CI is green. A squash or rebase merge writes a commit no reviewer saw, and CI fails it after the merge.
44
+ 7. Emit events the coordinator can wait on, one line each: `PROGRAM_EVENT <tranche> <PLAN|RED|COMMIT <sha>|WORKER_START <role> <model> <sha>|WORKER_END <role> <model> <verdict>|FALLBACK <role> <from>→<to> <reason>|PLAN_REVIEW <sha>|TEST_REVIEW <sha>|CODE_REVIEW <sha>|FINAL_REVIEW <sha>|PR_OPENED <url>|MERGED <sha>|BLOCKED <reason>|HEARTBEAT>`.
45
+
46
+ ```yaml
47
+ # <root>/<tranche>/evidence.yaml
48
+ tranche: api
49
+ status: open # ready once final_review approves
50
+ base_sha: 3f2a91c
51
+ approved_sha: c77d0a2 # required when ready
52
+ reviews:
53
+ - stage: code_review
54
+ sha: a4c0f11
55
+ author: { model: glm, run: glm-green-1 }
56
+ reviewer: { model: claude, run: claude-code-1 }
57
+ fallback: [{ role: reviewer, from: glm, kind: rate_limit, failures: 3, window: { from: "2026-09-25T09:00:00Z", to: "2026-09-25T09:08:00Z" }, reason: "429 from the provider on 3 attempts in 10 minutes" }]
58
+ verdict: request_changes
59
+ checked: [ACC-API-001, src/wagons/api, coder.bun.error-response-*]
60
+ findings:
61
+ - { id: F1, severity: high, evidence: "src/wagons/api/handler.ts:42", invariant: "coded error bodies", affects: [ui], proposed_fix: "return { code: 'API_NOT_FOUND' }" } # outcome added once a fresh code_review confirms the fix
62
+ report: delivery/api/code_review-1.json
63
+ ```
64
+
65
+ ## Default commands
66
+
67
+ `{prompt}` and `{worktree}` are substituted. Override per model under `delivery.commands.<model>.author` / `.review` when a CLI changes.
68
+
69
+ | Model | Author | Review (read-only) |
70
+ |---|---|---|
71
+ | glm | `zcode -p="{prompt}" --cwd {worktree} --mode edit` | `zcode -p="{prompt}" --cwd {worktree} --mode plan` |
72
+ | claude | `cd {worktree} && claude -p "{prompt}" --permission-mode acceptEdits --allowedTools "Bash(bun:*)" "Bash(git:*)" "Bash(gh pr:*)" --output-format json` | `cd {worktree} && claude -p "{prompt}" --allowedTools Read Grep Glob "Bash(git show:*)" "Bash(git diff:*)" "Bash(git log:*)" "Bash(bun test:*)" "Bash(bun run atdd-bun all)" "Bash(bun run atdd-bun planner)" "Bash(bun run atdd-bun tester)" "Bash(bun run atdd-bun coder security)" "Bash(bun run atdd-bun traceability)" "Bash(bun run atdd-bun delivery)" --disallowedTools Edit Write NotebookEdit --output-format json` |
73
+ | codex | `codex exec --cd {worktree} --sandbox workspace-write "{prompt}"` | `codex exec --cd {worktree} --sandbox read-only "{prompt}"` |
74
+
75
+ Before a hosted model receives private repository content, confirm the user or organization authorized it.
76
+
77
+ Never edit this skill, `review.md`, the conventions, or loosen the `delivery:` policy to get a tranche through. If the policy must change, stop and ask the human.
@@ -0,0 +1,45 @@
1
+ <!-- Generated by @afokapu/atdd-bun {{VERSION}}. Do not edit; run `atdd-bun agent init --replace`. -->
2
+ # Review contract
3
+
4
+ You are an independent reviewer for one stage of one tranche. The driver gave you the stage and the exact SHA.
5
+
6
+ **Read only.** You run in a detached worktree at the SHA under review. Do not edit, commit, or run anything that writes to it (including output-file options such as `git diff --output=`); the driver checks it is untouched afterwards. You may read files, search, use `git show`/`git diff`/`git log`, and run the gates (`bun test`, `bun run atdd-bun …`). You judge and propose; the author applies. If you edit, your review does not count.
7
+
8
+ ## How to review
9
+
10
+ Apply all three, every time:
11
+
12
+ - **Systematic.** Work through the stage's checklist below completely. List everything you checked in `checked`, not only what failed, so a skipped area is visible.
13
+ - **Systemic.** Look beyond the diff: contracts, other tranches, downstream owners, invariants the change could break elsewhere. Name them in each finding's `affects`.
14
+ - **Adversarial.** Assume the work is wrong and try to prove it: inputs that fail, tests that pass for the wrong reason, ways around a gate, rules satisfied in letter only. A finding needs concrete evidence (file:line, a failing command, a counter-example); "might be an issue" is not a finding.
15
+
16
+ ## Checklist by stage
17
+
18
+ - `plan_review`: the decomposition (wagon → WMBT → acceptance → train → journey → contract) covers the intent; every acceptance is testable and has one observable outcome; every WMBT has a SMOKE acceptance; dependencies and owned files do not overlap other tranches; `bun run atdd-bun planner` passes.
19
+ - `test_review`: every acceptance has a RED test bound to it by URN; each test fails for the missing behaviour and would still fail for a wrong implementation; no test asserts on mocks where the acceptance names observable output; `bun run atdd-bun tester` passes.
20
+ - `code_review`: every behaviour is correct against its acceptance, including edge and error paths; layering, composition, DTO and error-response rules hold; no security fault; SMOKE runs through the real entry point; `bun test` and `bun run atdd-bun all` pass.
21
+ - `final_review`: the whole PR at its head SHA, read as an architecture: coherent with the plan, no drift from what the earlier stages approved, no change a stage did not review, safe for every downstream owner.
22
+
23
+ ## Disputes
24
+
25
+ If a finding carries the author's `rebuttal`, judge it against the evidence. Withdraw it (leave it out of your findings) if the rebuttal holds. Uphold it (repeat it with the same `id`) if not. You settle it; there is no second round.
26
+
27
+ ## Output
28
+
29
+ Return exactly one YAML document, nothing else. The driver appends it to the tranche's evidence record.
30
+
31
+ ```yaml
32
+ stage: code_review # the stage you were given
33
+ sha: c77d0a2 # the SHA you reviewed
34
+ verdict: request_changes # or approve (approve only with no critical or high finding)
35
+ checked: [ACC-API-001, ACC-API-002, src/wagons/api, coder.bun.error-response-*]
36
+ findings:
37
+ - id: F1 # keep an upheld finding's id
38
+ severity: high # critical | high | medium | low
39
+ evidence: src/wagons/api/handler.ts:42 returns a bare string
40
+ invariant: every error response carries a coded body
41
+ affects: [ui]
42
+ proposed_fix: return { code "API_NOT_FOUND" } instead of the bare string
43
+ ```
44
+
45
+ Keep `proposed_fix` to a precise description or a short snippet, never a rewrite.
@@ -4,7 +4,7 @@ on:
4
4
  pull_request:
5
5
  merge_group:
6
6
  push:
7
- branches: [main, master]
7
+ branches: [{{PROTECTED_BRANCHES}}]
8
8
  jobs:
9
9
  atdd-bun:
10
10
  name: atdd-bun
@@ -21,6 +21,11 @@ jobs:
21
21
  - run: bun install --frozen-lockfile
22
22
  - run: bun run atdd-bun integrity
23
23
  - run: bun run atdd-bun all
24
+ env:
25
+ # The delivery profile's gate: before the merge on pull requests and the merge queue, after it on a push.
26
+ ATDD_DELIVERY_GATE: ${{ github.event_name == 'push' && 'post-merge' || 'merge' }}
27
+ # On a push, the gate judges everything since the previous tip, not only the last commit.
28
+ ATDD_BASE_REF: ${{ github.event_name == 'push' && github.event.before || '' }}
24
29
  - run: bun test
25
30
  - run: bun run atdd-bun release check
26
31
  - uses: actions/upload-artifact@v4