codex-orchestrator 0.1.23 → 0.1.25

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (40) hide show
  1. package/CHANGELOG.md +20 -0
  2. package/README.md +136 -282
  3. package/dist/src/config/schema.d.ts +26 -0
  4. package/dist/src/config/schema.d.ts.map +1 -1
  5. package/dist/src/config/schema.js +54 -0
  6. package/dist/src/config/schema.js.map +1 -1
  7. package/dist/src/runner/daemon-command.d.ts.map +1 -1
  8. package/dist/src/runner/daemon-command.js +25 -1
  9. package/dist/src/runner/daemon-command.js.map +1 -1
  10. package/dist/src/runner/durable-run-summary.d.ts +41 -0
  11. package/dist/src/runner/durable-run-summary.d.ts.map +1 -0
  12. package/dist/src/runner/durable-run-summary.js +66 -0
  13. package/dist/src/runner/durable-run-summary.js.map +1 -0
  14. package/dist/src/runner/fresh-context-review.d.ts +22 -0
  15. package/dist/src/runner/fresh-context-review.d.ts.map +1 -0
  16. package/dist/src/runner/fresh-context-review.js +157 -0
  17. package/dist/src/runner/fresh-context-review.js.map +1 -0
  18. package/dist/src/runner/handoff-evidence.d.ts +16 -0
  19. package/dist/src/runner/handoff-evidence.d.ts.map +1 -1
  20. package/dist/src/runner/handoff-evidence.js +62 -0
  21. package/dist/src/runner/handoff-evidence.js.map +1 -1
  22. package/dist/src/runner/plan-auto-command.d.ts.map +1 -1
  23. package/dist/src/runner/plan-auto-command.js +134 -28
  24. package/dist/src/runner/plan-auto-command.js.map +1 -1
  25. package/dist/src/runner/review-gate-policy.d.ts.map +1 -1
  26. package/dist/src/runner/review-gate-policy.js +14 -2
  27. package/dist/src/runner/review-gate-policy.js.map +1 -1
  28. package/dist/src/runner/rework-policy.d.ts +3 -0
  29. package/dist/src/runner/rework-policy.d.ts.map +1 -0
  30. package/dist/src/runner/rework-policy.js +20 -0
  31. package/dist/src/runner/rework-policy.js.map +1 -0
  32. package/dist/src/runner/scoped-auto-command.d.ts.map +1 -1
  33. package/dist/src/runner/scoped-auto-command.js +99 -27
  34. package/dist/src/runner/scoped-auto-command.js.map +1 -1
  35. package/dist/src/setup/project-config.d.ts.map +1 -1
  36. package/dist/src/setup/project-config.js +58 -0
  37. package/dist/src/setup/project-config.js.map +1 -1
  38. package/docs/deep-dive.md +364 -0
  39. package/package.json +2 -1
  40. package/prompts/workflows/scoped-implementation.md +9 -3
@@ -0,0 +1,364 @@
1
+ # codex-orchestrator Deep Dive
2
+
3
+ This document is the technical companion to the main README. The README explains
4
+ what problem `codex-orchestrator` solves and how to start using it. This file
5
+ describes how the package works internally, what features it provides, and where
6
+ the main control points are.
7
+
8
+ ## System Role
9
+
10
+ `codex-orchestrator` is a local runner that connects four things:
11
+
12
+ - GitHub Issues as the work queue;
13
+ - GitHub labels as the authorization and state model;
14
+ - git worktrees as isolated implementation workspaces;
15
+ - Codex CLI as the implementation agent.
16
+
17
+ The runner does not replace human review. It prepares work for review. A
18
+ successful run ends with a pushed branch, a draft pull request, issue comments,
19
+ and review labels. It does not auto-merge.
20
+
21
+ ## Package vs Repository Policy
22
+
23
+ The npm package contains reusable orchestration logic:
24
+
25
+ - CLI commands;
26
+ - config schema and validation;
27
+ - GitHub issue and PR operations;
28
+ - worktree and branch management;
29
+ - Codex prompt execution;
30
+ - validation and review-gate evaluation;
31
+ - durable local state and recovery reports.
32
+
33
+ Each target repository owns its policy under `.codex-orchestrator/`:
34
+
35
+ - `config.json` for labels, branches, checks, gates, deny rules, and runner
36
+ behavior;
37
+ - `prompts/` for repo-local prompts used by Codex workflows;
38
+ - proof and artifact directories created during runs.
39
+
40
+ This split keeps the package generic while letting each repository choose how
41
+ strict autonomous work should be.
42
+
43
+ ## CLI Commands
44
+
45
+ ### `health`
46
+
47
+ Checks whether the CLI can run in the current environment. It is intended as a
48
+ quick sanity check after installation.
49
+
50
+ ### `setup`
51
+
52
+ Creates project-local config and prompt files under `.codex-orchestrator/`.
53
+
54
+ Important behavior:
55
+
56
+ - infers GitHub owner and repo from `git remote origin` by default;
57
+ - can override owner, repo, or target path with flags;
58
+ - can create missing GitHub labels with `--prepare-labels`;
59
+ - can preview changes with `--dry-run`;
60
+ - copies package-owned fallback prompts when local skills are missing;
61
+ - does not launch Codex, commit files, or open pull requests.
62
+
63
+ ### `status`
64
+
65
+ Reads GitHub Issues and local runner state, then reports:
66
+
67
+ - issues eligible for autonomous work;
68
+ - skipped issues and the reason they were skipped;
69
+ - blocked or recoverable local state.
70
+
71
+ `status` is read-only. It does not mutate GitHub or launch Codex.
72
+
73
+ ### `run`
74
+
75
+ Executes one selected issue when its labels and state allow autonomous work.
76
+
77
+ `run` supports both authorization modes:
78
+
79
+ - `agent:auto` for one scoped implementation issue;
80
+ - `agent:plan-auto` for parent planning and issue-tree execution.
81
+
82
+ ### `daemon`
83
+
84
+ Polls GitHub Issues for eligible autonomous work and runs one issue at a time.
85
+
86
+ The daemon applies the configured issue selection policy after safety filters.
87
+ By default, priority labels sort eligible issues and issue number is the
88
+ deterministic tie-breaker.
89
+
90
+ After polling, the daemon can clean up runner-owned worktrees whose pull
91
+ requests have already been merged. Dirty, blocked, active, or unpublished
92
+ worktrees are preserved.
93
+
94
+ ## Issue Authorization Model
95
+
96
+ The runner only starts work that is explicitly authorized.
97
+
98
+ Default labels:
99
+
100
+ - `agent:auto` authorizes one scoped implementation run;
101
+ - `agent:plan-auto` authorizes parent planning and child issue execution;
102
+ - `agent:child` marks child issues that belong to an autonomous parent tree;
103
+ - `agent:running` means a runner has claimed the issue;
104
+ - `agent:blocked` means maintainer input or manual recovery is needed;
105
+ - `agent:manual` reserves the issue for human work;
106
+ - `agent:review` means the result is ready for human review.
107
+
108
+ Issues are skipped when they are closed, manual, blocked, already running,
109
+ already in review, or otherwise not authorized by policy.
110
+
111
+ Child issues are not inferred from ordinary GitHub links, milestones, project
112
+ fields, or casual references. They must carry the configured child label and the
113
+ runner-owned parent marker.
114
+
115
+ ## Scoped Issue Run
116
+
117
+ For an `agent:auto` issue, the runner follows this lifecycle:
118
+
119
+ 1. Load config and validate policy.
120
+ 2. Fetch the issue from GitHub.
121
+ 3. Check labels and state.
122
+ 4. Claim the issue with the running label.
123
+ 5. Create a runner-owned branch and git worktree.
124
+ 6. Build a prompt from the issue, config, and scoped implementation workflow.
125
+ 7. Run Codex CLI in the issue worktree.
126
+ 8. Read the Codex completion report.
127
+ 9. Collect the full local change set.
128
+ 10. Run configured validation checks.
129
+ 11. Run visual proof when required.
130
+ 12. Evaluate quality gates and deny rules.
131
+ 13. Optionally run bounded rework for machine-checkable blockers.
132
+ 14. Optionally run Fresh-Context Review.
133
+ 15. Write durable run evidence.
134
+ 16. Push the branch.
135
+ 17. Open a draft pull request.
136
+ 18. Post the review report and move the issue to review.
137
+
138
+ If a blocking condition is found, the runner does not publish the result. It
139
+ marks the issue blocked, preserves useful local evidence, and reports the
140
+ reason.
141
+
142
+ ## Parent Planning and Child Waves
143
+
144
+ `agent:plan-auto` is for larger work that should be planned before
145
+ implementation.
146
+
147
+ The parent flow can:
148
+
149
+ - ask Codex to produce or update the PRD;
150
+ - break the parent into child implementation issues;
151
+ - review the breakdown;
152
+ - triage child issues;
153
+ - identify which children are safe for autonomous execution;
154
+ - run children in dependency-aware waves;
155
+ - merge successful child branches into one integration branch;
156
+ - validate the integration branch;
157
+ - open one integration draft PR.
158
+
159
+ Parallel child execution is bounded by `runner.maxParallelChildren`. Child work
160
+ uses separate worktrees so concurrent runs do not share a mutable workspace.
161
+
162
+ ## Full Change-Set Awareness
163
+
164
+ The runner evaluates the whole result of an agent run:
165
+
166
+ - local commits;
167
+ - staged files;
168
+ - unstaged files;
169
+ - untracked files.
170
+
171
+ This matters because Codex may be allowed to create local commits. Those commits
172
+ are still treated as untrusted agent output until the runner validates them.
173
+
174
+ The runner owns external publication:
175
+
176
+ - pushing branches;
177
+ - opening draft PRs;
178
+ - moving labels;
179
+ - posting comments;
180
+ - merging child branches into integration branches;
181
+ - publishing packages or deploying.
182
+
183
+ If agent output attempts to bypass those boundaries, the run is blocked.
184
+
185
+ ## Validation Checks
186
+
187
+ Configured checks live in `.codex-orchestrator/config.json`.
188
+
189
+ Typical checks include commands such as:
190
+
191
+ ```json
192
+ {
193
+ "checks": {
194
+ "test": "npm test",
195
+ "typecheck": "npm run typecheck"
196
+ }
197
+ }
198
+ ```
199
+
200
+ By default, a missing `npm run <script>` command is reported as a skipped
201
+ warning instead of a hard failure. Repositories can change that with
202
+ `checksPolicy.missingNpmScript`.
203
+
204
+ For repositories with existing lint debt, `checksPolicy.lintBaseline.mode` can
205
+ be set to `touched-only`. In that mode, a repo-wide lint failure can be
206
+ downgraded when a separate touched-files lint command passes.
207
+
208
+ ## Review Gates
209
+
210
+ Review gates are runner-enforced checks that decide whether a result can be
211
+ published for human review.
212
+
213
+ The quality gate can require:
214
+
215
+ - strict TDD red-to-green evidence;
216
+ - a changed test file when runtime code changed;
217
+ - code review evidence;
218
+ - cleanup review evidence for larger runtime changes.
219
+
220
+ Runtime and test paths are configurable through:
221
+
222
+ - `reviewGates.quality.runtimeChangedPathGlobs`;
223
+ - `reviewGates.quality.testChangedPathGlobs`.
224
+
225
+ The visual proof gate can require screenshots or another runner-owned proof
226
+ command when issue text or changed paths indicate UI work.
227
+
228
+ ## Visual Proof
229
+
230
+ Visual proof is intentionally runner-owned. Codex can implement the UI, but the
231
+ runner decides whether proof is required and executes the configured proof
232
+ command after implementation.
233
+
234
+ The proof command usually runs browser automation, such as Playwright. The
235
+ runner provides environment variables for:
236
+
237
+ - issue number;
238
+ - artifact directory;
239
+ - proof directory;
240
+ - Playwright profile directory;
241
+ - worktree path;
242
+ - changed files.
243
+
244
+ Screenshot files created under the proof directory are attached to the PR and
245
+ issue review report. If a proof command exits successfully but does not create
246
+ the configured minimum number of screenshots, the runner reports a warning.
247
+
248
+ For Android UI work, the implementation prompt uses device-backed proof through
249
+ `adb` or an emulator. Missing Android tooling or no usable device is reported as
250
+ a concrete warning instead of a release blocker by itself.
251
+
252
+ ## Loop Policy
253
+
254
+ Loop Policy controls runner-owned automation around retries and evidence.
255
+
256
+ It includes:
257
+
258
+ - issue selection priority labels and tie-breaker;
259
+ - bounded rework attempts;
260
+ - retryable blocker types;
261
+ - Fresh-Context Review;
262
+ - Durable Run Summaries;
263
+ - non-mutating Policy Suggestions.
264
+
265
+ Bounded rework is limited to machine-checkable blockers such as missing or
266
+ invalid completion reports, no changed files, failed configured checks, or
267
+ missing quality-gate evidence. It stops at the configured attempt limit.
268
+
269
+ Fresh-Context Review runs a separate Codex session with the issue, diff, and
270
+ validation evidence. It does not reuse the implementation transcript. In the
271
+ current config model, the mode is advisory; repositories can choose whether
272
+ high-confidence policy violations block publication.
273
+
274
+ Durable Run Summaries record the outcome, confirmed facts, validation, blockers,
275
+ residual risks, next action, and policy suggestions. They reference existing
276
+ logs and reports; they do not replace them.
277
+
278
+ Policy Suggestions are report-only. They never edit prompts, config, labels, or
279
+ issue state.
280
+
281
+ ## Deny Rules
282
+
283
+ Deny rules block publication when the agent result touches forbidden areas or
284
+ attempts unsafe actions.
285
+
286
+ The default policy can block:
287
+
288
+ - secret files;
289
+ - destructive database or cache actions;
290
+ - production deploy or release actions;
291
+ - additional repository-defined path globs.
292
+
293
+ These rules are evaluated before a result is published.
294
+
295
+ ## Durable State and Recovery
296
+
297
+ The runner keeps local state so interrupted work can be inspected and recovered.
298
+
299
+ Durable evidence can include:
300
+
301
+ - agent output;
302
+ - completion reports;
303
+ - validation results;
304
+ - skipped checks;
305
+ - blocked reasons;
306
+ - visual artifacts;
307
+ - run summaries;
308
+ - preserved worktrees.
309
+
310
+ The runner preserves worktrees when deleting them would hide useful evidence,
311
+ for example when they are dirty, blocked, active, or unpublished.
312
+
313
+ ## Prompt and Workflow System
314
+
315
+ Workflows are configured in `config.json`.
316
+
317
+ The default workflow set includes:
318
+
319
+ - PRD creation or update;
320
+ - issue breakdown;
321
+ - breakdown review;
322
+ - triage;
323
+ - scoped implementation;
324
+ - issue-tree orchestration.
325
+
326
+ Each workflow points to either a package-owned fallback prompt or a compatible
327
+ local skill/prompt copied during setup. This lets the package work out of the
328
+ box while still allowing repositories to customize agent behavior.
329
+
330
+ ## Config Surface
331
+
332
+ The top-level config areas are:
333
+
334
+ - `github` for owner, repo, label preparation, and label definitions;
335
+ - `runner` for workspace root, child concurrency, state directory, local commit
336
+ policy, and worktree cleanup;
337
+ - `codex` for Codex CLI command, args, timeouts, and prompt/report env vars;
338
+ - `project` for config and prompt directories;
339
+ - `workflows` for prompt or skill routing;
340
+ - `checks` and `checksPolicy` for validation commands;
341
+ - `reviewGates` for quality and visual proof requirements;
342
+ - `loopPolicy` for issue selection, rework, review, summaries, and suggestions;
343
+ - `deny` for secret and unsafe-action protection;
344
+ - `branches` for branch templates;
345
+ - `pullRequests` for PR title templates;
346
+ - `issueClassification` for promotion criteria and clarification behavior.
347
+
348
+ Runtime state must not be committed as config. The schema rejects known runtime
349
+ keys in committed config files.
350
+
351
+ ## Current Boundaries
352
+
353
+ The package currently focuses on local runner workflows:
354
+
355
+ - explicit one-off runs;
356
+ - daemon polling;
357
+ - project-local config;
358
+ - runner-owned worktree cleanup;
359
+ - GitHub Issues and Pull Requests;
360
+ - Codex CLI as the agent backend.
361
+
362
+ Hosted infrastructure, non-GitHub issue trackers, and non-Codex agents are not
363
+ part of the current package, although the code keeps adapter boundaries for
364
+ future expansion.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "codex-orchestrator",
3
- "version": "0.1.23",
3
+ "version": "0.1.25",
4
4
  "description": "Reusable GitHub Issues runner for Codex.",
5
5
  "type": "module",
6
6
  "license": "MIT",
@@ -22,6 +22,7 @@
22
22
  "files": [
23
23
  "dist/src",
24
24
  "prompts",
25
+ "docs/deep-dive.md",
25
26
  "CHANGELOG.md",
26
27
  "README.md",
27
28
  "LICENSE"
@@ -11,6 +11,12 @@ Run code-review before completion for runtime changes and report the result.
11
11
  Report validation, skipped checks, and risks.
12
12
  For UI or visual changes, follow the orchestration prompt's visual proof
13
13
  contract. If a runner-owned visual proof command is configured, prepare its
14
- script/artifacts (prefer Playwright) but let the runner execute it. If visual
15
- proof is not possible in this environment, state that explicitly in skipped
16
- checks with the concrete reason and proceed.
14
+ script/artifacts (prefer Playwright for browser/web UI) but let the runner
15
+ execute it. For Android mobile app UI work, use device-backed proof instead of
16
+ Playwright: run `adb devices -l`, prefer a connected non-emulator device serial,
17
+ and run `export ANDROID_SERIAL=<serial>`. Otherwise run `emulator -list-avds`,
18
+ start an AVD in a separate shell with `emulator -avd <avd-name>`, and wait with
19
+ `adb wait-for-device`. If Test Android Apps skills are unavailable, try to enable
20
+ or load that plugin through the available Codex plugin/tool discovery mechanism.
21
+ If the plugin cannot be enabled, or no usable device or emulator is available,
22
+ report that as a warning/skipped check with the concrete reason and proceed.