codex-orchestrator 2.0.3 → 2.0.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (161) hide show
  1. package/CHANGELOG.md +51 -427
  2. package/README.md +161 -37
  3. package/dist/src/index.d.ts +1 -1
  4. package/dist/src/index.d.ts.map +1 -1
  5. package/dist/src/v2/acceptance-proof.d.ts +5 -0
  6. package/dist/src/v2/acceptance-proof.d.ts.map +1 -1
  7. package/dist/src/v2/acceptance-proof.js +10 -2
  8. package/dist/src/v2/acceptance-proof.js.map +1 -1
  9. package/dist/src/v2/adapters/gh-issue-adapter.d.ts.map +1 -1
  10. package/dist/src/v2/adapters/gh-issue-adapter.js +6 -7
  11. package/dist/src/v2/adapters/gh-issue-adapter.js.map +1 -1
  12. package/dist/src/v2/adapters/gh-pull-request-adapter.d.ts +7 -1
  13. package/dist/src/v2/adapters/gh-pull-request-adapter.d.ts.map +1 -1
  14. package/dist/src/v2/adapters/gh-pull-request-adapter.js +288 -0
  15. package/dist/src/v2/adapters/gh-pull-request-adapter.js.map +1 -1
  16. package/dist/src/v2/adapters/pull-requests.d.ts +69 -0
  17. package/dist/src/v2/adapters/pull-requests.d.ts.map +1 -1
  18. package/dist/src/v2/adapters/pull-requests.js +48 -0
  19. package/dist/src/v2/adapters/pull-requests.js.map +1 -1
  20. package/dist/src/v2/adapters/worktree.d.ts +1 -0
  21. package/dist/src/v2/adapters/worktree.d.ts.map +1 -1
  22. package/dist/src/v2/adapters/worktree.js +10 -1
  23. package/dist/src/v2/adapters/worktree.js.map +1 -1
  24. package/dist/src/v2/cli-contract.d.ts +3 -3
  25. package/dist/src/v2/cli-contract.d.ts.map +1 -1
  26. package/dist/src/v2/cli-contract.js +1 -3
  27. package/dist/src/v2/cli-contract.js.map +1 -1
  28. package/dist/src/v2/cli.d.ts +33 -0
  29. package/dist/src/v2/cli.d.ts.map +1 -0
  30. package/dist/src/v2/{candidate-cli.js → cli.js} +55 -39
  31. package/dist/src/v2/cli.js.map +1 -0
  32. package/dist/src/v2/code-review-report.d.ts +1 -1
  33. package/dist/src/v2/code-review-report.d.ts.map +1 -1
  34. package/dist/src/v2/code-review-report.js +2 -2
  35. package/dist/src/v2/code-review-report.js.map +1 -1
  36. package/dist/src/v2/codex-process.d.ts.map +1 -1
  37. package/dist/src/v2/codex-process.js +12 -1
  38. package/dist/src/v2/codex-process.js.map +1 -1
  39. package/dist/src/v2/config.d.ts +2 -3
  40. package/dist/src/v2/config.d.ts.map +1 -1
  41. package/dist/src/v2/config.js +0 -3
  42. package/dist/src/v2/config.js.map +1 -1
  43. package/dist/src/v2/contained-report-operation.d.ts +2 -2
  44. package/dist/src/v2/contained-report-operation.d.ts.map +1 -1
  45. package/dist/src/v2/contained-report-operation.js +1 -1
  46. package/dist/src/v2/contained-report-operation.js.map +1 -1
  47. package/dist/src/v2/containment.d.ts +15 -2
  48. package/dist/src/v2/containment.d.ts.map +1 -1
  49. package/dist/src/v2/containment.js +43 -6
  50. package/dist/src/v2/containment.js.map +1 -1
  51. package/dist/src/v2/direct-delivery.d.ts +5 -10
  52. package/dist/src/v2/direct-delivery.d.ts.map +1 -1
  53. package/dist/src/v2/direct-delivery.js +32 -91
  54. package/dist/src/v2/direct-delivery.js.map +1 -1
  55. package/dist/src/v2/proof-report.d.ts.map +1 -1
  56. package/dist/src/v2/proof-report.js +55 -29
  57. package/dist/src/v2/proof-report.js.map +1 -1
  58. package/dist/src/v2/review-feedback-coordinator.d.ts +54 -0
  59. package/dist/src/v2/review-feedback-coordinator.d.ts.map +1 -0
  60. package/dist/src/v2/review-feedback-coordinator.js +245 -0
  61. package/dist/src/v2/review-feedback-coordinator.js.map +1 -0
  62. package/dist/src/v2/review-feedback.d.ts +127 -0
  63. package/dist/src/v2/review-feedback.d.ts.map +1 -0
  64. package/dist/src/v2/review-feedback.js +436 -0
  65. package/dist/src/v2/review-feedback.js.map +1 -0
  66. package/dist/src/v2/run-issue.d.ts +63 -9
  67. package/dist/src/v2/run-issue.d.ts.map +1 -1
  68. package/dist/src/v2/run-issue.js +798 -78
  69. package/dist/src/v2/run-issue.js.map +1 -1
  70. package/dist/src/v2/run-store.d.ts +49 -5
  71. package/dist/src/v2/run-store.d.ts.map +1 -1
  72. package/dist/src/v2/run-store.js +138 -44
  73. package/dist/src/v2/run-store.js.map +1 -1
  74. package/dist/src/v2/runtime.d.ts +40 -3
  75. package/dist/src/v2/runtime.d.ts.map +1 -1
  76. package/dist/src/v2/runtime.js +245 -52
  77. package/dist/src/v2/runtime.js.map +1 -1
  78. package/dist/src/v2/setup-cli.d.ts.map +1 -1
  79. package/dist/src/v2/setup-cli.js +4 -11
  80. package/dist/src/v2/setup-cli.js.map +1 -1
  81. package/dist/src/v2/setup-runtime.d.ts.map +1 -1
  82. package/dist/src/v2/setup-runtime.js +1 -61
  83. package/dist/src/v2/setup-runtime.js.map +1 -1
  84. package/dist/src/v2/setup-store.d.ts +0 -5
  85. package/dist/src/v2/setup-store.d.ts.map +1 -1
  86. package/dist/src/v2/setup-store.js +3 -106
  87. package/dist/src/v2/setup-store.js.map +1 -1
  88. package/dist/src/v2/setup.d.ts +6 -46
  89. package/dist/src/v2/setup.d.ts.map +1 -1
  90. package/dist/src/v2/setup.js +12 -294
  91. package/dist/src/v2/setup.js.map +1 -1
  92. package/dist/src/v2/workflow-assets.d.ts +19 -11
  93. package/dist/src/v2/workflow-assets.d.ts.map +1 -1
  94. package/dist/src/v2/workflow-assets.js +132 -40
  95. package/dist/src/v2/workflow-assets.js.map +1 -1
  96. package/docs/deep-dive.md +328 -56
  97. package/internal-workflow/docs/agents/bugfix-quality-gate.md +11 -0
  98. package/internal-workflow/docs/agents/coding-skill-routing.md +116 -196
  99. package/internal-workflow/docs/agents/contract-test-ledger.md +11 -1
  100. package/internal-workflow/docs/agents/review-gates.md +32 -39
  101. package/internal-workflow/docs/agents/review-protocol.md +75 -147
  102. package/internal-workflow/evals/coding-skill-evals.json +84 -0
  103. package/internal-workflow/manifest.json +1 -1
  104. package/internal-workflow/operations/acceptance-proof/SKILL.md +7 -1
  105. package/internal-workflow/operations/ambiguity-review/SKILL.md +2 -0
  106. package/internal-workflow/operations/code-review/SKILL.md +21 -1
  107. package/internal-workflow/operations/implementation/SKILL.md +22 -1
  108. package/internal-workflow/operations/spec-author/SKILL.md +10 -1
  109. package/internal-workflow/operations/spec-review/SKILL.md +10 -1
  110. package/internal-workflow/operations/triage/SKILL.md +10 -1
  111. package/internal-workflow/schemas/code-review-v1.json +1 -1
  112. package/internal-workflow/schemas/proof-report-v1.json +1 -1
  113. package/internal-workflow/skills/agent-auto/SKILL.md +6 -1
  114. package/internal-workflow/skills/code-debugger/SKILL.md +122 -0
  115. package/internal-workflow/skills/code-debugger/agents/openai.yaml +7 -0
  116. package/internal-workflow/skills/code-review/SKILL.md +51 -17
  117. package/internal-workflow/skills/code-review/references/cleanup-lens.md +52 -0
  118. package/internal-workflow/skills/implementation-spec-maker/SKILL.md +15 -6
  119. package/internal-workflow/skills/implementation-spec-maker/agents/openai.yaml +1 -1
  120. package/internal-workflow/skills/implementation-spec-maker/references/spec-template.md +2 -2
  121. package/internal-workflow/skills/implementation-spec-review/SKILL.md +108 -204
  122. package/internal-workflow/skills/implementation-spec-review/evals/evals.json +24 -0
  123. package/internal-workflow/skills/implementation-spec-review/references/review-loop.md +93 -0
  124. package/internal-workflow/skills/small-task-implementer/SKILL.md +15 -8
  125. package/internal-workflow/skills/spec-implementer/SKILL.md +101 -172
  126. package/internal-workflow/skills/spec-implementer/evals/evals.json +30 -0
  127. package/internal-workflow/skills/spec-implementer/references/review-loop.md +100 -0
  128. package/internal-workflow/skills/tdd/SKILL.md +20 -6
  129. package/internal-workflow/skills/tdd/agents/openai.yaml +2 -2
  130. package/internal-workflow/skills/tdd/evals/evals.json +18 -0
  131. package/internal-workflow/skills/tdd/mocking.md +3 -42
  132. package/internal-workflow/skills/tdd/refactoring.md +6 -8
  133. package/package.json +9 -6
  134. package/dist/src/v2/adapters/target-activity-fence.d.ts +0 -23
  135. package/dist/src/v2/adapters/target-activity-fence.d.ts.map +0 -1
  136. package/dist/src/v2/adapters/target-activity-fence.js +0 -249
  137. package/dist/src/v2/adapters/target-activity-fence.js.map +0 -1
  138. package/dist/src/v2/candidate-cli.d.ts +0 -26
  139. package/dist/src/v2/candidate-cli.d.ts.map +0 -1
  140. package/dist/src/v2/candidate-cli.js.map +0 -1
  141. package/dist/src/v2/legacy-cutover.d.ts +0 -52
  142. package/dist/src/v2/legacy-cutover.d.ts.map +0 -1
  143. package/dist/src/v2/legacy-cutover.js +0 -87
  144. package/dist/src/v2/legacy-cutover.js.map +0 -1
  145. package/internal-workflow/docs/agents/artifact-review-loop.md +0 -267
  146. package/internal-workflow/docs/agents/implementation-review-loop.md +0 -302
  147. package/internal-workflow/operations/cleanup-review/SKILL.md +0 -3
  148. package/internal-workflow/operations/spec-implementation/SKILL.md +0 -3
  149. package/internal-workflow/profiles/implementer_deep.toml +0 -9
  150. package/internal-workflow/profiles/researcher_standard.toml +0 -9
  151. package/internal-workflow/profiles/reviewer_fast.toml +0 -9
  152. package/internal-workflow/skills/cleanup-review/SKILL.md +0 -84
  153. package/internal-workflow/skills/cleanup-review/agents/openai.yaml +0 -6
  154. package/internal-workflow/skills/codebase-design/DEEPENING.md +0 -35
  155. package/internal-workflow/skills/codebase-design/DESIGN-IT-TWICE.md +0 -50
  156. package/internal-workflow/skills/codebase-design/SKILL.md +0 -82
  157. package/internal-workflow/skills/codebase-design/agents/openai.yaml +0 -6
  158. package/internal-workflow/skills/research/SKILL.md +0 -107
  159. package/internal-workflow/skills/research/agents/openai.yaml +0 -6
  160. package/internal-workflow/skills/ui-evidence-proof/SKILL.md +0 -123
  161. package/internal-workflow/skills/ui-evidence-proof/agents/openai.yaml +0 -6
package/docs/deep-dive.md CHANGED
@@ -1,93 +1,365 @@
1
- # Runner architecture
1
+ # Runner architecture and execution model
2
2
 
3
- ## One public runtime
3
+ This document describes the current package runtime as implemented under `src/v2/`. It is the technical companion to the user-oriented README: the README explains how to operate the package; this document explains ownership, state transitions, worker contracts, validation, recovery, and publication.
4
4
 
5
- The package bin is `dist/src/v2/candidate-cli.js`; `candidate-cli` is only the historical source filename. It is the sole public CLI and routes `setup`, `doctor`, `status`, direct `run`, and serial `daemon` commands. The root package export exposes only V2 contracts.
5
+ ## 1. System boundary
6
6
 
7
- The runtime is split into a policy core under `src/v2/` and a small package-owned adapter closure under `src/v2/adapters/`. No earlier runner implementation is shipped or executed.
7
+ Codex Orchestrator has one public runtime and one delivery authority.
8
8
 
9
- `internal-workflow/manifest.json` is the sole workflow inventory. Release-time sync imports an explicit skill/profile/doc closure and assigns each operation one entrypoint, schema, profile, file closure, and authority policy. At runtime the Runner verifies the full package tree, publishes one immutable generation with a no-replace receipt, pins that receipt in the run record, and creates operation-scoped attempt snapshots only from the pin. Package updates and conflicting user skills cannot change an active run.
9
+ The npm binary is `dist/src/v2/cli.js`. It is the only public command entrypoint and exposes:
10
10
 
11
- ## Trust boundary
11
+ - `setup`
12
+ - `doctor`
13
+ - `status`
14
+ - `run`
15
+ - `daemon`
12
16
 
13
- The Runner is trusted. It owns:
17
+ Direct `run` and daemon-discovered issues both enter the same `RunIssue.runIssue` lifecycle. The daemon only adds serial polling around that lifecycle. The package root export exposes V2 contracts, and the published tarball contains only the compiled V2 closure, the generated internal workflow, public documentation, changelog, license, and package metadata.
14
18
 
15
- - issue discovery and authorization;
16
- - worktree and durable state ownership;
17
- - finite check execution;
18
- - process launch, timeout, cancellation, and quiescence;
19
- - Git commits, branches, pushes, pull requests, labels, and comments;
20
- - mobile leases and proof artifact validation.
19
+ Reusable orchestration policy lives in `src/v2/`. Git, GitHub, process, durable-file, browser, and mobile integrations live in `src/v2/adapters/` or other package-owned V2 runtime modules. Target-specific policy lives in `.codex-orchestrator/config.json`; it is not inferred from worker prompts.
21
20
 
22
- Implementation and proof Codex processes are untrusted workers. They receive bounded input, a dedicated tool home, a safe `PATH`, and only explicitly allowed non-secret tool environment values. Ordinary Codex execution and native Codex subagents may use the same user-owned Codex authentication and can read files available to the same local OS user; the containment certificate records this explicitly. GitHub, SSH, npm, and cloud publication credentials are scrubbed, and operations that require external authority remain finite Runner-owned actions.
21
+ ## 2. Trusted Runner and untrusted workers
23
22
 
24
- This boundary prevents prompt content or repository code from receiving external publication authority. It does not claim OS-level secrecy from the current local user, so reports and artifacts are separately checked for credential/path disclosure. A child can implement, inspect, test, and report; the Runner performs authorized publication after validation.
23
+ The trusted Runner owns every action that can authorize work, change durable orchestration state, or publish externally:
25
24
 
26
- ## Configuration
25
+ - repository identity and configuration validation;
26
+ - issue discovery and `agent:auto` authorization;
27
+ - repository ownership locks and fencing;
28
+ - run records, worktrees, branch identity, and Git snapshots;
29
+ - workflow generation materialization and verification;
30
+ - worker launch, timeout, cancellation, and process-group quiescence;
31
+ - configured checks;
32
+ - proof capabilities and artifact validation;
33
+ - commits, pushes, draft pull requests, labels, and issue comments;
34
+ - browser/mobile proof policy and device leases.
27
35
 
28
- `.codex-orchestrator/config.json` is an exact, versioned contract. Unknown keys and removed policy surfaces are rejected. It names one GitHub repository, five labels (`auto`, `running`, `blocked`, `review`, `waitingHuman`), one branch template, a serial polling interval, five implementation cycles, one Codex command contract, finite checks, one proof artifact directory, and deny lists.
36
+ Codex processes are operation-scoped workers. The current internal workflow defines seven operations: `triage`, `ambiguity-review`, `spec-author`, `spec-review`, `implementation`, `code-review`, and `acceptance-proof`. A worker receives a dedicated tool home, a restricted environment, a safe `PATH`, bounded prompt facts, a report schema, and only the files appropriate to its operation.
29
37
 
30
- The committed config contains policy, never live run state or credentials. Durable runtime state is stored beneath the configured state directory and the Runner's private home.
38
+ Workers do not receive GitHub, SSH, npm, or cloud publication credentials. They cannot turn a proposed shell command into Runner authority and cannot directly push, create a pull request, alter issue labels, or publish comments.
31
39
 
32
- ## `runIssue`
40
+ This is an authority-containment boundary, not an OS sandbox. Ordinary Codex execution and native Codex subagents use the same local OS account and may use the user's existing Codex authentication or read files available to that user. The containment certificate records that accepted local-read risk and binds the installed `codex` command's reported version, canonical executable path and digest, plus the orchestrator package version; configuration does not pin a Codex release. Environment scrubbing, denied paths, isolated read views, and report/artifact validation prevent those local capabilities from becoming an external publication grant.
33
41
 
34
- Direct runs and daemon-discovered runs call the same lifecycle:
42
+ ## 3. Strict repository configuration
35
43
 
36
- 1. Snapshot issue and repository identity.
37
- 2. Verify authorization and acquire fenced ownership.
38
- 3. Create or reconcile the issue worktree.
39
- 4. Persist a reviewed route. An `awaiting-user` route publishes one durable question, freezes only a current WRITE+ answer, restores running labels, and reruns triage in the same Run before implementation.
40
- 5. Run one implementation attempt and validate its structured report.
41
- 6. Inspect all tracked, staged, unstaged, untracked, and ignored denied-path changes.
42
- 7. Commit the validated implementation candidate locally.
43
- 8. Run configured checks and create a nominal `CheckedChange` bound to exact Git and content hashes.
44
- 9. Run Acceptance Proof against that binding.
45
- 10. Publish with durable intents and postcondition reconciliation.
46
- 11. Persist and return one typed terminal or resumable result.
44
+ `.codex-orchestrator/config.json` is an exact versioned contract with schema `codex-orchestrator.agent-auto`, version `2`. Unknown, missing, removed, or malformed keys are rejected.
47
45
 
48
- Implementation findings return to the same worktree for at most five cycles. A malformed report gets one report-only repair; a clean transport disconnect gets one separate transport retry. Neither budget silently becomes another implementation cycle.
46
+ The config contains:
49
47
 
50
- ## Durable effects
48
+ - `github`: canonical owner/repository, base branch, and five distinct label definitions;
49
+ - `runner`: worktree root, durable state directory, fixed branch template, daemon polling interval, and the five-cycle limit;
50
+ - `codex`: command, total and idle timeouts, and denied worker tool network;
51
+ - `checks`: a finite map of check IDs to Runner-owned commands;
52
+ - `proof.artifactDir`: proof-owned artifact root inside the issue worktree;
53
+ - `deny.readPaths`: repository-relative or canonical absolute paths protected from workers;
54
+ - `deny.commands`: canonical absolute command paths excluded from worker authority.
51
55
 
52
- Every external or non-idempotent operation uses intent-before-effect and confirmation-after-observation. On restart, the Runner inspects the durable intent and remote/local postcondition before deciding whether to retry. It does not infer failure merely from a lost response.
56
+ Default setup creates:
53
57
 
54
- Atomic files use write, flush, rename, and directory synchronization where supported. Locks and leases carry fencing tokens and process/boot identity. Unknown ownership or inability to prove process-group absence is a safe halt, not permission to continue.
58
+ ```text
59
+ .codex-orchestrator/config.json
60
+ .codex-orchestrator/workspaces-v2/
61
+ .codex-orchestrator/v2/state/
62
+ .codex-orchestrator/v2/proofs/
63
+ ```
55
64
 
56
- Workflow generation and attempt publication use hard-link claims, recovery chains, sealed file evidence, and ready receipts. Existing content is reused only after full path, mode, owner, size, and hash verification; active pre-generation V1 state fails closed while terminal V1 history remains readable.
65
+ The three runtime directories are placed in a managed `.gitignore` block. Setup adds `npm test` and `npm run typecheck` to `checks` only when matching scripts exist in the target `package.json`.
57
66
 
58
- ## Checks and publication
67
+ `setup`, `doctor`, and `status` use the same strict parser. Setup can create or verify the current policy and optionally prepare labels; doctor/status are read-only. Older or unrelated config shapes are not execution authority and are rejected instead of loading a compatibility runtime.
59
68
 
60
- Checks are configured finite commands executed by the Runner. Arbitrary agent-proposed shell commands do not become Runner authority. A check result is bound into `CheckedChange`; any later Git, index, tracked-content, untracked-content, worktree, or check-policy drift invalidates proof.
69
+ ## 4. Immutable package-owned workflow
61
70
 
62
- Only the Runner publishes. The implementation and proof processes cannot push, open a pull request, alter labels, or post comments. Publication is resumable and verifies exact repository, branch, commit, and remote postconditions.
71
+ `internal-workflow/manifest.json` is the sole inventory of worker operations. Maintainers regenerate it with `npm run refresh:workflow` from the explicit source declaration in `scripts/agent-auto-workflow-source.json` and the allowlisted files under `${CODEX_HOME:-$HOME/.codex}`.
63
72
 
64
- ## Acceptance Proof
73
+ The generated inventory binds each operation to:
65
74
 
66
- Acceptance Proof receives the issue snapshot, frozen criteria, and nominal checked-change capability. It runs in a separate contained Codex process and writes only proof-owned artifacts.
75
+ - an operation ID and profile;
76
+ - its primary and dependency skills;
77
+ - shared routing resources;
78
+ - an exact file closure;
79
+ - input/output schemas and wrapper resources;
80
+ - its authority policy;
81
+ - package-owned eval suites used by maintainers.
82
+
83
+ Compilation rejects missing dependencies, stale bytes, undeclared resources, or adapters that fail to link all declared authorities. Eval files are included in the immutable generation and schema-validated, but excluded from operation snapshots so workers cannot consume expected or forbidden test answers.
84
+
85
+ At the beginning of a new run, the Runner verifies the packaged workflow tree and materializes one immutable generation under the private orchestrator home. Publication uses no-replace claims, sealed file evidence, hashes, and a ready receipt. The generation receipt and skill hashes are persisted in the run record. Every later operation is created from that pin, so a package update or conflicting consumer skill cannot alter an active run.
86
+
87
+ `npm run check:workflow` compares local workflow sources with committed generated bytes without writing. `npm run verify:workflow` verifies the committed package without consulting local or consumer skills.
88
+
89
+ ## 5. Issue eligibility and ownership
90
+
91
+ For a new run, the Runner performs these checks before allowing worker execution:
92
+
93
+ 1. Parse the strict config and derive the lowercase canonical repository identity.
94
+ 2. Acquire the repository owner lock. A known live owner returns `requeued`; ambiguous ownership returns a resumable safety block.
95
+ 3. Re-read the config after lock acquisition and require byte-identical policy and repository identity.
96
+ 4. Validate the current containment certificate.
97
+ 5. Read the issue and durable run state.
98
+ 6. Require an open issue with `agent:auto`, without running, blocked, review, or waiting-human labels.
99
+ 7. Require that no open pull request already owns `codex/issue-${issueNumber}` against the configured base branch.
100
+
101
+ For an eligible issue, the Runner freezes the issue snapshot and acceptance criteria, resolves the base SHA, creates the run record, and persists an intent to claim labels before changing GitHub. The issue is moved to the exact running label set and receives one marker-bound claim comment. Only then does the Runner create `.codex-orchestrator/workspaces-v2/issue-${issueNumber}` on `codex/issue-${issueNumber}`.
102
+
103
+ Every later externally meaningful phase revalidates authorization. The open issue must still carry `agent:auto` and `agent:running`, and exactly one trusted claim comment must bind the run ID, issue, and branch. Revoked or conflicting authority fails closed.
104
+
105
+ ## 6. Routing before implementation
106
+
107
+ After claim and worktree creation, the Runner invokes `triage` against the frozen issue facts, acceptance criteria, repository, base SHA, and pinned workflow generation. Triage must return a schema-valid, evidence-backed route:
108
+
109
+ - `direct`
110
+ - `spec-required`
111
+ - `awaiting-user`
112
+ - `blocked`
113
+
114
+ The route report records inspected evidence, explicit assumptions, and route-specific details. The Runner hashes the report and persists a route receipt bound to the workflow generation.
115
+
116
+ A malformed triage report has one report repair budget. A clean transport failure has one separate transport retry. These retries do not become implementation cycles.
117
+
118
+ An `awaiting-user` proposal is privileged because it pauses autonomous work. It must describe at least two materially different observable product outcomes and prove that repository authority does not select between them. A separate `ambiguity-review` worker receives the candidate and either approves or rejects it. Only one candidate review is allowed. A rejected candidate can use the single triage repair path; an approved candidate becomes the durable route receipt.
119
+
120
+ The route determines the downstream lifecycle:
121
+
122
+ ```mermaid
123
+ flowchart TD
124
+ A["Eligible issue claimed"] --> B["Triage"]
125
+ B -->|"direct"| C["Implementation and delivery loop"]
126
+ B -->|"spec-required"| D["Spec author and independent review"]
127
+ B -->|"awaiting-user"| E["Approved question and trusted answer"]
128
+ B -->|"blocked"| F["Typed terminal blocker"]
129
+ D --> G["Frozen specification receipt"]
130
+ E --> B
131
+ C --> H["Draft PR and review-ready"]
132
+ ```
133
+
134
+ ### Direct route
135
+
136
+ The issue already contains enough behavioral authority and verification detail for deterministic implementation. The Runner enters the bounded implementation/review/check/proof loop described below.
137
+
138
+ ### Specification-required route
139
+
140
+ Complexity, cross-cutting behavior, or insufficient executable detail makes direct implementation unsafe even though no product decision is missing. The Runner starts a durable specification state machine:
141
+
142
+ 1. `spec-author` creates a revision in `author` mode.
143
+ 2. `spec-review` independently reviews the exact revision in `full` mode.
144
+ 3. A needs-work verdict preserves a defect ledger and returns to `spec-author` in repair mode.
145
+ 4. The same reviewer session performs closure review against the affected defects.
146
+ 5. Only an approved revision with resolved blockers is frozen.
147
+
148
+ Prepared and launched invocation records are persisted before and after process launch. Recovery proves process-group absence before replacing an uncertain invocation. Malformed reports and transport retries are bounded; exhaustion becomes a typed blocker rather than an unbounded author/reviewer loop.
149
+
150
+ The terminal result for this route is `spec-frozen` with an immutable `FrozenSpecReceipt`. The current runtime does not silently continue from a newly authored specification into implementation. That boundary keeps specification approval separately auditable.
151
+
152
+ ### Awaiting-user route
153
+
154
+ After independent ambiguity approval, the waiting-human coordinator:
155
+
156
+ 1. creates a marker-bound question receipt;
157
+ 2. persists comment intent before posting the question;
158
+ 3. changes labels from running to `agent:auto` + `agent:waiting-human`;
159
+ 4. scans only post-question replies using the required answer prefix;
160
+ 5. accepts answers only from a current repository writer with sufficient permission;
161
+ 6. normalizes equivalent answers and freezes their hashes;
162
+ 7. treats conflicting trusted answers as a bounded clarification, not an arbitrary choice;
163
+ 8. revalidates answer authority before restoring running labels;
164
+ 9. archives the waiting episode and reruns triage in the same run with the trusted answer.
165
+
166
+ The first unanswered pass returns `awaiting-user`. At most one follow-up question is allowed; unresolved conflict or repeated ambiguity exhausts the waiting budget. A frozen answer is immutable and cannot be replaced by a later edited comment.
167
+
168
+ ## 7. Direct implementation and independent review
169
+
170
+ The direct route uses one issue worktree for at most five implementation cycles. Each cycle is bound to the same run, frozen issue criteria, route receipt, and workflow generation.
171
+
172
+ An `implementation` worker receives the current cycle and any findings from prior review, checks, or proof. It must return a structured implementation report. The Runner then:
173
+
174
+ - verifies that denied paths did not change;
175
+ - validates the exact report schema;
176
+ - allows one report-only repair when the report is malformed;
177
+ - proves that report repair did not change the worktree;
178
+ - distinguishes an external blocker from a completed implementation;
179
+ - requires the branch head to remain at the frozen base SHA;
180
+ - inventories tracked, staged, unstaged, untracked, and denied-path state;
181
+ - requires the reported changed-file list to equal the observed change set.
182
+
183
+ A clean implementation transport failure receives one separate retry only if the complete Git freshness baseline is unchanged. Any unexplained mutation converts that retry into a safety block.
184
+
185
+ Before configured checks, a separate `code-review` worker reviews a fingerprint of the complete implementation target. The fingerprint binds Git freshness, changed files, route decision, workflow generation, cycle, and frozen criteria. Full review covers at least acceptance criteria, correctness, and test quality; cleanup is a lens within this final review.
186
+
187
+ Review maintains an append-preserving defect ledger. If review returns `needs-work`, open defects become implementation findings and consume the next implementation cycle. After repair, the same reviewer session performs closure review only for affected defects while preserving unrelated and previously accepted findings. Approved review requires every blocker or execution risk to be verified or explicitly superseded.
188
+
189
+ Review invocation intent, process IDs, report hashes, transport retries, report repairs, target revisions, and target fingerprints are durable. A crash after launch cannot cause a replacement review until process absence is proven. Review target drift after approval is a safety failure.
190
+
191
+ ## 8. Configured checks and `CheckedChange`
192
+
193
+ After review clears, the Runner executes the finite commands in `config.checks`. Workers cannot add commands to this set. Each result records its ID, exact command, pass/fail status, and output hash.
194
+
195
+ A failed check becomes a durable repair finding and starts another implementation cycle if the five-cycle budget remains. Passed checks are reused on a safe resume, but any new repair cycle clears stale check and proof bindings.
196
+
197
+ Once all checks pass, the Runner fingerprints the reviewed files and content, stages the complete validated change, and verifies that staging did not change either binding. It then mints a `CheckedChange` capability containing:
198
+
199
+ - canonical repository, run ID, issue number, and cycle;
200
+ - base and head SHA;
201
+ - index tree SHA;
202
+ - tracked and untracked content hashes;
203
+ - worktree identity;
204
+ - exact changed files;
205
+ - passed check records and check-policy hash;
206
+ - package and proof schema versions.
207
+
208
+ Any later change to Git, the index, tracked or untracked content, worktree identity, changed files, or check policy invalidates the capability.
209
+
210
+ ## 9. Acceptance Proof
211
+
212
+ Acceptance Proof is independent from both implementation and code review. The `acceptance-proof` worker receives the frozen issue criteria and a nominal checked-change capability. It runs in a separate contained process and may write only below its proof-owned artifact root.
67
213
 
68
214
  The Runner validates:
69
215
 
70
- - exact report schema and criterion IDs;
71
- - artifact root containment, size, hash, UTF-8, and freshness;
72
- - credentials in every text artifact;
73
- - host identity and publication type for public artifacts;
216
+ - the exact Proof Report schema and status semantics;
217
+ - one result for every frozen criterion ID;
218
+ - evidence references and required confidence;
219
+ - artifact root containment and canonical paths;
220
+ - regular-file type, size, hash, UTF-8 validity, and freshness;
221
+ - credential absence in all text artifacts;
222
+ - public artifact type and host-identity removal;
74
223
  - no product diff during proof;
75
- - browser workflow and viewport evidence when visual;
76
- - exact Android or iOS lease ownership for mobile evidence;
77
- - unchanged checked-change freshness after proof.
224
+ - current browser workflow and viewport evidence for browser targets;
225
+ - exact Android or iOS lease ownership for mobile targets;
226
+ - checked-change freshness again after proof.
227
+
228
+ Local command output and static-inspection evidence may contain machine paths because they are never public. Public evidence is restricted to screenshots or sanitized generated summaries under the stricter publication contract. Credentials are forbidden in both local and public evidence.
229
+
230
+ Browser proof validates current workflow evidence rather than accepting an isolated screenshot. Mobile proof uses Runner-owned leases and refuses to take over a user-owned device, app, IDE, or Flutter session.
231
+
232
+ `needs-rework` findings return to the same implementation loop and consume another cycle. External, safety, malformed, quiescence, or exhausted outcomes are mapped to typed run results. Only `passed` proof produces a proof receipt and permits publication.
233
+
234
+ ## 10. Runner-owned publication
235
+
236
+ Workers never publish. After proof passes, the Runner performs publication in this order:
237
+
238
+ 1. Revalidate issue authorization.
239
+ 2. Persist commit intent bound to parent SHA, tree SHA, and message.
240
+ 3. Create or reconcile the single implementation commit.
241
+ 4. Persist push intent and verify the remote branch SHA.
242
+ 5. Persist pull-request intent and find or create the marker-bound draft PR.
243
+ 6. Verify the PR head, base, marker, and repository identity.
244
+ 7. Persist and publish the final issue comment.
245
+ 8. Replace running labels with the exact review-ready label set.
246
+ 9. Write local terminal evidence and persist `review-ready` with the PR URL.
247
+
248
+ Every step checks its postcondition before proceeding. A conflicting remote branch, duplicate marker, unexpected PR, changed tree, revoked authorization, or ambiguous effect becomes a safety or transport result; it is never treated as implicit success.
249
+
250
+ ### Post-PR review continuation
251
+
252
+ A successful direct run persists review-feedback state inside the same atomic
253
+ run record. Version 1 state remains readable; the next compare-and-swap emits
254
+ version 2. The first observation of a migrated `review-ready` run baselines
255
+ already-present eligible source IDs without launching a worker or changing
256
+ GitHub, so an upgrade cannot retrospectively execute old feedback.
257
+
258
+ The daemon discovers both `agent:auto` and `agent:review`, deduplicates issue
259
+ numbers, and sends every candidate through `RunIssue.runIssue`. Unchanged
260
+ `review-ready` output is suppressed only in daemon memory; durable run state
261
+ remains authoritative. One-shot diagnostics and live smoke may additionally
262
+ constrain daemon execution to one discovered issue with `--once --issue`; the
263
+ ordinary long-running daemon always processes the complete discovered set. A V2 idle run observes one coherent PR snapshot bounded
264
+ by equal identity/head reads around all GraphQL thread and REST review pages.
265
+ Eligible inputs are:
266
+
267
+ - the non-empty root of an unresolved, non-outdated inline thread bound to the
268
+ observed head; or
269
+ - a non-empty submitted `CHANGES_REQUESTED` review bound to that head.
270
+
271
+ Each exposed body requires a fresh repository `write` or `admin` permission
272
+ receipt tied to the immutable author ID. Replies can contribute IDs and hashes
273
+ to drift detection, but their bodies are never persisted or sent to a worker.
274
+ Bots, ordinary PR conversation comments, approvals, blank or dismissed reviews,
275
+ read-only identities, and consumed source IDs are excluded.
276
+
277
+ Activation is one atomic transition from outer `review-ready` to
278
+ `implementing`: it freezes and consumes the batch, reserves feedback round one,
279
+ opens the existing direct-review repair ledger with `pr-review` provenance,
280
+ clears terminal/check/CheckedChange/proof receipts, and persists the exact
281
+ `agent:auto` + `agent:running` label intent. Feedback rounds are independently
282
+ bounded to 1–3; completed needs-work results alone reserve the next round. The
283
+ original five-cycle counter is unchanged.
284
+
285
+ Every worker launch and publication effect revalidates PR identity, refs,
286
+ marker, source content, immutable author, and current permission. `pre-update`
287
+ requires the old published head and exact thread state. After the fast-forward
288
+ push, `post-push` requires the new head and unchanged trusted source, while
289
+ allowing only GitHub-derived resolved/outdated changes. Safety or exhaustion
290
+ blocks are non-resumable.
291
+
292
+ Terminal feedback safety/exhaustion cleanup performs only a monotonic authority
293
+ reduction from the exact run-owned running or review label set to
294
+ `agent:auto` + `agent:blocked`. This cleanup remains allowed when the claim was
295
+ just removed, because it cannot launch work or publish product state.
296
+
297
+ Update publication has separate durable commit, push, PR-summary, and final-label
298
+ intents. Local HEAD and the remote branch must begin at the persisted published
299
+ head; the new commit must be its single parent-child successor. Unknown delivery
300
+ is adopted only when the exact parent, tree, message, branch, and resulting SHA
301
+ match. No reset, rebase, amend, or force push exists. Success posts one
302
+ `review-feedback:<batch>` PR summary, returns the same issue and PR to
303
+ `agent:review`/`review-ready`, and records the new published head. GitHub review
304
+ threads are never auto-resolved.
305
+
306
+ ## 11. Durable state and crash recovery
307
+
308
+ The durable run file is stored beneath the configured state directory. Each run record contains identity, lifecycle, cycle budgets, frozen issue data, workflow pin, route state, waiting-human history, spec/review and review-feedback state, process records, checks, proof bindings, publication intent, and terminal outcome.
309
+
310
+ State updates use generation-based compare-and-swap. Atomic files use write, flush, rename, and directory synchronization where supported. Locks and leases include fencing plus process and boot identity.
311
+
312
+ External and non-idempotent effects follow intent-before-effect and confirmation-after-observation:
313
+
314
+ ```text
315
+ persist intent -> perform finite effect -> observe exact postcondition -> clear intent
316
+ ```
317
+
318
+ If the process exits after the effect but before confirmation, the next invocation reads the intent and checks the local or remote postcondition. It does not infer failure from a lost response and does not blindly repeat the effect.
319
+
320
+ Prepared worker invocations can be abandoned without launch. Launched invocations require positive process-group absence before replacement. If quiescence cannot be proven, the run enters `safe-halt` or a non-resumable transport/safety outcome. Unknown process state is never permission to relaunch against the same worktree.
321
+
322
+ An existing nonterminal issue run is resumed only when its canonical repository, branch, worktree path, base SHA, workflow generation, authorization, and lifecycle invariants still match. Terminal outcomes are replayed without re-executing the workflow, except that a direct `review-ready` checkpoint may perform the bounded trusted-feedback observation described above.
323
+
324
+ ## 12. Result and failure model
325
+
326
+ The CLI prints exact JSON envelopes. Run results include:
327
+
328
+ - `review-ready`: direct delivery completed and the draft PR is verified;
329
+ - `spec-frozen`: the specification route completed with an immutable receipt;
330
+ - `awaiting-user`: a durable question is waiting for an authorized answer;
331
+ - `not-eligible`: the issue was never claimed because public eligibility failed;
332
+ - `requeued`: a known live owner currently holds the repository;
333
+ - `blocked`: a typed `external`, `safety`, or `exhausted` condition;
334
+ - `transport-failed`: effect delivery or observation could not be safely confirmed;
335
+ - `cancelled`: cancellation was observed and persisted;
336
+ - `internal-error`: an invariant, schema, or local operation failed outside a safe domain result.
337
+
338
+ `blocked.resumable` and `transport-failed.resumable` are part of the contract. A resumable result still requires the external condition to be corrected; it does not bypass reconciliation on the next call. `evidencePath` identifies the durable local evidence record for the result.
339
+
340
+ Exit codes are grouped for automation:
78
341
 
79
- Command output and static inspection may remain local and include machine paths. They are not public evidence. Screenshots and sanitized generated summaries may be publishable when their stricter checks pass.
342
+ - `0`: successful or intentionally paused progress such as `review-ready`, `spec-frozen`, `awaiting-user`, or `requeued`;
343
+ - `20`: blocked policy outcome;
344
+ - `21`: not eligible;
345
+ - `70`: transport or internal failure;
346
+ - `130`: cancelled.
80
347
 
81
- ## Setup and cutover
348
+ Daemon mode processes discovered issues serially and returns the greatest observed exit severity for the polling pass.
82
349
 
83
- `setup` creates or verifies V2 config. `--prepare-labels` performs only the requested GitHub label preparation. `doctor` and `status` are read-only inspections.
350
+ ## 13. Package verification and live validation
84
351
 
85
- Exact Config V1 is upgraded atomically by `setup configure` only after the shared run-owner fence and remote running claims are proven absent. Operational commands return `migration-required` until that upgrade completes.
352
+ Local package verification is:
86
353
 
87
- Recognized earlier config is parsed only by the bounded cutover reader. `setup --fresh` acquires both ownership fences, proves no active old claims, saves immutable backup evidence, and publishes V2 config last. The old runtime never executes as part of this process.
354
+ ```sh
355
+ npm run refresh:workflow
356
+ npm run typecheck
357
+ npm test
358
+ npm pack --dry-run --json
359
+ ```
88
360
 
89
- ## Live validation
361
+ The build deletes `dist` before TypeScript compilation so removed modules cannot survive in tests or the tarball. `prepack` verifies the committed workflow and rebuilds from a clean output directory.
90
362
 
91
- Local validation is `npm run typecheck`, `npm test`, and `npm pack --dry-run --json`. Build removes `dist` first so deleted modules cannot leak into tests or tarballs.
363
+ `npm run smoke:live` packs and installs the exact package bytes into a temporary consumer and mutates only the configured scratch GitHub repository. The default `core-release` profile proves package installation through real model-backed operations, browser evidence, and a safety-negative path. Cleanup verifies that run-owned issues, PRs, branches, labels, worktrees, and temporary directories are absent.
92
364
 
93
- The default live smoke packs and installs the exact candidate bytes in a temporary consumer and uses a scratch GitHub repository. Its compact release profile proves package installation, one normal default Codex run, browser evidence, and a safety-negative path. Cleanup verifies that run-owned issues, pull requests, branches, labels, and temporary directories are absent.
365
+ Live smoke is not a normal local test and must run only with explicit authorization. Release publication is owned by the GitHub release workflow after the release commit reaches `main`.
@@ -0,0 +1,11 @@
1
+ # Bugfix Quality Gate
2
+
3
+ For bug fixes, do not claim completion until these are clear:
4
+
5
+ - Invariant: final user/system-visible outcome that must be true.
6
+ - Boundary: paths, states, async events, retries, caches, workers, or integrations that can affect it.
7
+ - Proof: test/log/smoke/check that fails on the old behavior and verifies the final outcome, not only an intermediate signal.
8
+ - Negative proof: what the proof does not cover.
9
+ - Claim limit: final response must not claim broader coverage than the changed code and proof support.
10
+
11
+ If the bug spans state, async, lifecycle, retries, cache, auth, persistence, or cross-module contracts, include the competing condition in the regression proof when feasible.