codex-orchestrator 0.1.23 → 0.1.25
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +20 -0
- package/README.md +136 -282
- package/dist/src/config/schema.d.ts +26 -0
- package/dist/src/config/schema.d.ts.map +1 -1
- package/dist/src/config/schema.js +54 -0
- package/dist/src/config/schema.js.map +1 -1
- package/dist/src/runner/daemon-command.d.ts.map +1 -1
- package/dist/src/runner/daemon-command.js +25 -1
- package/dist/src/runner/daemon-command.js.map +1 -1
- package/dist/src/runner/durable-run-summary.d.ts +41 -0
- package/dist/src/runner/durable-run-summary.d.ts.map +1 -0
- package/dist/src/runner/durable-run-summary.js +66 -0
- package/dist/src/runner/durable-run-summary.js.map +1 -0
- package/dist/src/runner/fresh-context-review.d.ts +22 -0
- package/dist/src/runner/fresh-context-review.d.ts.map +1 -0
- package/dist/src/runner/fresh-context-review.js +157 -0
- package/dist/src/runner/fresh-context-review.js.map +1 -0
- package/dist/src/runner/handoff-evidence.d.ts +16 -0
- package/dist/src/runner/handoff-evidence.d.ts.map +1 -1
- package/dist/src/runner/handoff-evidence.js +62 -0
- package/dist/src/runner/handoff-evidence.js.map +1 -1
- package/dist/src/runner/plan-auto-command.d.ts.map +1 -1
- package/dist/src/runner/plan-auto-command.js +134 -28
- package/dist/src/runner/plan-auto-command.js.map +1 -1
- package/dist/src/runner/review-gate-policy.d.ts.map +1 -1
- package/dist/src/runner/review-gate-policy.js +14 -2
- package/dist/src/runner/review-gate-policy.js.map +1 -1
- package/dist/src/runner/rework-policy.d.ts +3 -0
- package/dist/src/runner/rework-policy.d.ts.map +1 -0
- package/dist/src/runner/rework-policy.js +20 -0
- package/dist/src/runner/rework-policy.js.map +1 -0
- package/dist/src/runner/scoped-auto-command.d.ts.map +1 -1
- package/dist/src/runner/scoped-auto-command.js +99 -27
- package/dist/src/runner/scoped-auto-command.js.map +1 -1
- package/dist/src/setup/project-config.d.ts.map +1 -1
- package/dist/src/setup/project-config.js +58 -0
- package/dist/src/setup/project-config.js.map +1 -1
- package/docs/deep-dive.md +364 -0
- package/package.json +2 -1
- package/prompts/workflows/scoped-implementation.md +9 -3
|
@@ -0,0 +1,364 @@
|
|
|
1
|
+
# codex-orchestrator Deep Dive
|
|
2
|
+
|
|
3
|
+
This document is the technical companion to the main README. The README explains
|
|
4
|
+
what problem `codex-orchestrator` solves and how to start using it. This file
|
|
5
|
+
describes how the package works internally, what features it provides, and where
|
|
6
|
+
the main control points are.
|
|
7
|
+
|
|
8
|
+
## System Role
|
|
9
|
+
|
|
10
|
+
`codex-orchestrator` is a local runner that connects four things:
|
|
11
|
+
|
|
12
|
+
- GitHub Issues as the work queue;
|
|
13
|
+
- GitHub labels as the authorization and state model;
|
|
14
|
+
- git worktrees as isolated implementation workspaces;
|
|
15
|
+
- Codex CLI as the implementation agent.
|
|
16
|
+
|
|
17
|
+
The runner does not replace human review. It prepares work for review. A
|
|
18
|
+
successful run ends with a pushed branch, a draft pull request, issue comments,
|
|
19
|
+
and review labels. It does not auto-merge.
|
|
20
|
+
|
|
21
|
+
## Package vs Repository Policy
|
|
22
|
+
|
|
23
|
+
The npm package contains reusable orchestration logic:
|
|
24
|
+
|
|
25
|
+
- CLI commands;
|
|
26
|
+
- config schema and validation;
|
|
27
|
+
- GitHub issue and PR operations;
|
|
28
|
+
- worktree and branch management;
|
|
29
|
+
- Codex prompt execution;
|
|
30
|
+
- validation and review-gate evaluation;
|
|
31
|
+
- durable local state and recovery reports.
|
|
32
|
+
|
|
33
|
+
Each target repository owns its policy under `.codex-orchestrator/`:
|
|
34
|
+
|
|
35
|
+
- `config.json` for labels, branches, checks, gates, deny rules, and runner
|
|
36
|
+
behavior;
|
|
37
|
+
- `prompts/` for repo-local prompts used by Codex workflows;
|
|
38
|
+
- proof and artifact directories created during runs.
|
|
39
|
+
|
|
40
|
+
This split keeps the package generic while letting each repository choose how
|
|
41
|
+
strict autonomous work should be.
|
|
42
|
+
|
|
43
|
+
## CLI Commands
|
|
44
|
+
|
|
45
|
+
### `health`
|
|
46
|
+
|
|
47
|
+
Checks whether the CLI can run in the current environment. It is intended as a
|
|
48
|
+
quick sanity check after installation.
|
|
49
|
+
|
|
50
|
+
### `setup`
|
|
51
|
+
|
|
52
|
+
Creates project-local config and prompt files under `.codex-orchestrator/`.
|
|
53
|
+
|
|
54
|
+
Important behavior:
|
|
55
|
+
|
|
56
|
+
- infers GitHub owner and repo from `git remote origin` by default;
|
|
57
|
+
- can override owner, repo, or target path with flags;
|
|
58
|
+
- can create missing GitHub labels with `--prepare-labels`;
|
|
59
|
+
- can preview changes with `--dry-run`;
|
|
60
|
+
- copies package-owned fallback prompts when local skills are missing;
|
|
61
|
+
- does not launch Codex, commit files, or open pull requests.
|
|
62
|
+
|
|
63
|
+
### `status`
|
|
64
|
+
|
|
65
|
+
Reads GitHub Issues and local runner state, then reports:
|
|
66
|
+
|
|
67
|
+
- issues eligible for autonomous work;
|
|
68
|
+
- skipped issues and the reason they were skipped;
|
|
69
|
+
- blocked or recoverable local state.
|
|
70
|
+
|
|
71
|
+
`status` is read-only. It does not mutate GitHub or launch Codex.
|
|
72
|
+
|
|
73
|
+
### `run`
|
|
74
|
+
|
|
75
|
+
Executes one selected issue when its labels and state allow autonomous work.
|
|
76
|
+
|
|
77
|
+
`run` supports both authorization modes:
|
|
78
|
+
|
|
79
|
+
- `agent:auto` for one scoped implementation issue;
|
|
80
|
+
- `agent:plan-auto` for parent planning and issue-tree execution.
|
|
81
|
+
|
|
82
|
+
### `daemon`
|
|
83
|
+
|
|
84
|
+
Polls GitHub Issues for eligible autonomous work and runs one issue at a time.
|
|
85
|
+
|
|
86
|
+
The daemon applies the configured issue selection policy after safety filters.
|
|
87
|
+
By default, priority labels sort eligible issues and issue number is the
|
|
88
|
+
deterministic tie-breaker.
|
|
89
|
+
|
|
90
|
+
After polling, the daemon can clean up runner-owned worktrees whose pull
|
|
91
|
+
requests have already been merged. Dirty, blocked, active, or unpublished
|
|
92
|
+
worktrees are preserved.
|
|
93
|
+
|
|
94
|
+
## Issue Authorization Model
|
|
95
|
+
|
|
96
|
+
The runner only starts work that is explicitly authorized.
|
|
97
|
+
|
|
98
|
+
Default labels:
|
|
99
|
+
|
|
100
|
+
- `agent:auto` authorizes one scoped implementation run;
|
|
101
|
+
- `agent:plan-auto` authorizes parent planning and child issue execution;
|
|
102
|
+
- `agent:child` marks child issues that belong to an autonomous parent tree;
|
|
103
|
+
- `agent:running` means a runner has claimed the issue;
|
|
104
|
+
- `agent:blocked` means maintainer input or manual recovery is needed;
|
|
105
|
+
- `agent:manual` reserves the issue for human work;
|
|
106
|
+
- `agent:review` means the result is ready for human review.
|
|
107
|
+
|
|
108
|
+
Issues are skipped when they are closed, manual, blocked, already running,
|
|
109
|
+
already in review, or otherwise not authorized by policy.
|
|
110
|
+
|
|
111
|
+
Child issues are not inferred from ordinary GitHub links, milestones, project
|
|
112
|
+
fields, or casual references. They must carry the configured child label and the
|
|
113
|
+
runner-owned parent marker.
|
|
114
|
+
|
|
115
|
+
## Scoped Issue Run
|
|
116
|
+
|
|
117
|
+
For an `agent:auto` issue, the runner follows this lifecycle:
|
|
118
|
+
|
|
119
|
+
1. Load config and validate policy.
|
|
120
|
+
2. Fetch the issue from GitHub.
|
|
121
|
+
3. Check labels and state.
|
|
122
|
+
4. Claim the issue with the running label.
|
|
123
|
+
5. Create a runner-owned branch and git worktree.
|
|
124
|
+
6. Build a prompt from the issue, config, and scoped implementation workflow.
|
|
125
|
+
7. Run Codex CLI in the issue worktree.
|
|
126
|
+
8. Read the Codex completion report.
|
|
127
|
+
9. Collect the full local change set.
|
|
128
|
+
10. Run configured validation checks.
|
|
129
|
+
11. Run visual proof when required.
|
|
130
|
+
12. Evaluate quality gates and deny rules.
|
|
131
|
+
13. Optionally run bounded rework for machine-checkable blockers.
|
|
132
|
+
14. Optionally run Fresh-Context Review.
|
|
133
|
+
15. Write durable run evidence.
|
|
134
|
+
16. Push the branch.
|
|
135
|
+
17. Open a draft pull request.
|
|
136
|
+
18. Post the review report and move the issue to review.
|
|
137
|
+
|
|
138
|
+
If a blocking condition is found, the runner does not publish the result. It
|
|
139
|
+
marks the issue blocked, preserves useful local evidence, and reports the
|
|
140
|
+
reason.
|
|
141
|
+
|
|
142
|
+
## Parent Planning and Child Waves
|
|
143
|
+
|
|
144
|
+
`agent:plan-auto` is for larger work that should be planned before
|
|
145
|
+
implementation.
|
|
146
|
+
|
|
147
|
+
The parent flow can:
|
|
148
|
+
|
|
149
|
+
- ask Codex to produce or update the PRD;
|
|
150
|
+
- break the parent into child implementation issues;
|
|
151
|
+
- review the breakdown;
|
|
152
|
+
- triage child issues;
|
|
153
|
+
- identify which children are safe for autonomous execution;
|
|
154
|
+
- run children in dependency-aware waves;
|
|
155
|
+
- merge successful child branches into one integration branch;
|
|
156
|
+
- validate the integration branch;
|
|
157
|
+
- open one integration draft PR.
|
|
158
|
+
|
|
159
|
+
Parallel child execution is bounded by `runner.maxParallelChildren`. Child work
|
|
160
|
+
uses separate worktrees so concurrent runs do not share a mutable workspace.
|
|
161
|
+
|
|
162
|
+
## Full Change-Set Awareness
|
|
163
|
+
|
|
164
|
+
The runner evaluates the whole result of an agent run:
|
|
165
|
+
|
|
166
|
+
- local commits;
|
|
167
|
+
- staged files;
|
|
168
|
+
- unstaged files;
|
|
169
|
+
- untracked files.
|
|
170
|
+
|
|
171
|
+
This matters because Codex may be allowed to create local commits. Those commits
|
|
172
|
+
are still treated as untrusted agent output until the runner validates them.
|
|
173
|
+
|
|
174
|
+
The runner owns external publication:
|
|
175
|
+
|
|
176
|
+
- pushing branches;
|
|
177
|
+
- opening draft PRs;
|
|
178
|
+
- moving labels;
|
|
179
|
+
- posting comments;
|
|
180
|
+
- merging child branches into integration branches;
|
|
181
|
+
- publishing packages or deploying.
|
|
182
|
+
|
|
183
|
+
If agent output attempts to bypass those boundaries, the run is blocked.
|
|
184
|
+
|
|
185
|
+
## Validation Checks
|
|
186
|
+
|
|
187
|
+
Configured checks live in `.codex-orchestrator/config.json`.
|
|
188
|
+
|
|
189
|
+
Typical checks include commands such as:
|
|
190
|
+
|
|
191
|
+
```json
|
|
192
|
+
{
|
|
193
|
+
"checks": {
|
|
194
|
+
"test": "npm test",
|
|
195
|
+
"typecheck": "npm run typecheck"
|
|
196
|
+
}
|
|
197
|
+
}
|
|
198
|
+
```
|
|
199
|
+
|
|
200
|
+
By default, a missing `npm run <script>` command is reported as a skipped
|
|
201
|
+
warning instead of a hard failure. Repositories can change that with
|
|
202
|
+
`checksPolicy.missingNpmScript`.
|
|
203
|
+
|
|
204
|
+
For repositories with existing lint debt, `checksPolicy.lintBaseline.mode` can
|
|
205
|
+
be set to `touched-only`. In that mode, a repo-wide lint failure can be
|
|
206
|
+
downgraded when a separate touched-files lint command passes.
|
|
207
|
+
|
|
208
|
+
## Review Gates
|
|
209
|
+
|
|
210
|
+
Review gates are runner-enforced checks that decide whether a result can be
|
|
211
|
+
published for human review.
|
|
212
|
+
|
|
213
|
+
The quality gate can require:
|
|
214
|
+
|
|
215
|
+
- strict TDD red-to-green evidence;
|
|
216
|
+
- a changed test file when runtime code changed;
|
|
217
|
+
- code review evidence;
|
|
218
|
+
- cleanup review evidence for larger runtime changes.
|
|
219
|
+
|
|
220
|
+
Runtime and test paths are configurable through:
|
|
221
|
+
|
|
222
|
+
- `reviewGates.quality.runtimeChangedPathGlobs`;
|
|
223
|
+
- `reviewGates.quality.testChangedPathGlobs`.
|
|
224
|
+
|
|
225
|
+
The visual proof gate can require screenshots or another runner-owned proof
|
|
226
|
+
command when issue text or changed paths indicate UI work.
|
|
227
|
+
|
|
228
|
+
## Visual Proof
|
|
229
|
+
|
|
230
|
+
Visual proof is intentionally runner-owned. Codex can implement the UI, but the
|
|
231
|
+
runner decides whether proof is required and executes the configured proof
|
|
232
|
+
command after implementation.
|
|
233
|
+
|
|
234
|
+
The proof command usually runs browser automation, such as Playwright. The
|
|
235
|
+
runner provides environment variables for:
|
|
236
|
+
|
|
237
|
+
- issue number;
|
|
238
|
+
- artifact directory;
|
|
239
|
+
- proof directory;
|
|
240
|
+
- Playwright profile directory;
|
|
241
|
+
- worktree path;
|
|
242
|
+
- changed files.
|
|
243
|
+
|
|
244
|
+
Screenshot files created under the proof directory are attached to the PR and
|
|
245
|
+
issue review report. If a proof command exits successfully but does not create
|
|
246
|
+
the configured minimum number of screenshots, the runner reports a warning.
|
|
247
|
+
|
|
248
|
+
For Android UI work, the implementation prompt uses device-backed proof through
|
|
249
|
+
`adb` or an emulator. Missing Android tooling or no usable device is reported as
|
|
250
|
+
a concrete warning instead of a release blocker by itself.
|
|
251
|
+
|
|
252
|
+
## Loop Policy
|
|
253
|
+
|
|
254
|
+
Loop Policy controls runner-owned automation around retries and evidence.
|
|
255
|
+
|
|
256
|
+
It includes:
|
|
257
|
+
|
|
258
|
+
- issue selection priority labels and tie-breaker;
|
|
259
|
+
- bounded rework attempts;
|
|
260
|
+
- retryable blocker types;
|
|
261
|
+
- Fresh-Context Review;
|
|
262
|
+
- Durable Run Summaries;
|
|
263
|
+
- non-mutating Policy Suggestions.
|
|
264
|
+
|
|
265
|
+
Bounded rework is limited to machine-checkable blockers such as missing or
|
|
266
|
+
invalid completion reports, no changed files, failed configured checks, or
|
|
267
|
+
missing quality-gate evidence. It stops at the configured attempt limit.
|
|
268
|
+
|
|
269
|
+
Fresh-Context Review runs a separate Codex session with the issue, diff, and
|
|
270
|
+
validation evidence. It does not reuse the implementation transcript. In the
|
|
271
|
+
current config model, the mode is advisory; repositories can choose whether
|
|
272
|
+
high-confidence policy violations block publication.
|
|
273
|
+
|
|
274
|
+
Durable Run Summaries record the outcome, confirmed facts, validation, blockers,
|
|
275
|
+
residual risks, next action, and policy suggestions. They reference existing
|
|
276
|
+
logs and reports; they do not replace them.
|
|
277
|
+
|
|
278
|
+
Policy Suggestions are report-only. They never edit prompts, config, labels, or
|
|
279
|
+
issue state.
|
|
280
|
+
|
|
281
|
+
## Deny Rules
|
|
282
|
+
|
|
283
|
+
Deny rules block publication when the agent result touches forbidden areas or
|
|
284
|
+
attempts unsafe actions.
|
|
285
|
+
|
|
286
|
+
The default policy can block:
|
|
287
|
+
|
|
288
|
+
- secret files;
|
|
289
|
+
- destructive database or cache actions;
|
|
290
|
+
- production deploy or release actions;
|
|
291
|
+
- additional repository-defined path globs.
|
|
292
|
+
|
|
293
|
+
These rules are evaluated before a result is published.
|
|
294
|
+
|
|
295
|
+
## Durable State and Recovery
|
|
296
|
+
|
|
297
|
+
The runner keeps local state so interrupted work can be inspected and recovered.
|
|
298
|
+
|
|
299
|
+
Durable evidence can include:
|
|
300
|
+
|
|
301
|
+
- agent output;
|
|
302
|
+
- completion reports;
|
|
303
|
+
- validation results;
|
|
304
|
+
- skipped checks;
|
|
305
|
+
- blocked reasons;
|
|
306
|
+
- visual artifacts;
|
|
307
|
+
- run summaries;
|
|
308
|
+
- preserved worktrees.
|
|
309
|
+
|
|
310
|
+
The runner preserves worktrees when deleting them would hide useful evidence,
|
|
311
|
+
for example when they are dirty, blocked, active, or unpublished.
|
|
312
|
+
|
|
313
|
+
## Prompt and Workflow System
|
|
314
|
+
|
|
315
|
+
Workflows are configured in `config.json`.
|
|
316
|
+
|
|
317
|
+
The default workflow set includes:
|
|
318
|
+
|
|
319
|
+
- PRD creation or update;
|
|
320
|
+
- issue breakdown;
|
|
321
|
+
- breakdown review;
|
|
322
|
+
- triage;
|
|
323
|
+
- scoped implementation;
|
|
324
|
+
- issue-tree orchestration.
|
|
325
|
+
|
|
326
|
+
Each workflow points to either a package-owned fallback prompt or a compatible
|
|
327
|
+
local skill/prompt copied during setup. This lets the package work out of the
|
|
328
|
+
box while still allowing repositories to customize agent behavior.
|
|
329
|
+
|
|
330
|
+
## Config Surface
|
|
331
|
+
|
|
332
|
+
The top-level config areas are:
|
|
333
|
+
|
|
334
|
+
- `github` for owner, repo, label preparation, and label definitions;
|
|
335
|
+
- `runner` for workspace root, child concurrency, state directory, local commit
|
|
336
|
+
policy, and worktree cleanup;
|
|
337
|
+
- `codex` for Codex CLI command, args, timeouts, and prompt/report env vars;
|
|
338
|
+
- `project` for config and prompt directories;
|
|
339
|
+
- `workflows` for prompt or skill routing;
|
|
340
|
+
- `checks` and `checksPolicy` for validation commands;
|
|
341
|
+
- `reviewGates` for quality and visual proof requirements;
|
|
342
|
+
- `loopPolicy` for issue selection, rework, review, summaries, and suggestions;
|
|
343
|
+
- `deny` for secret and unsafe-action protection;
|
|
344
|
+
- `branches` for branch templates;
|
|
345
|
+
- `pullRequests` for PR title templates;
|
|
346
|
+
- `issueClassification` for promotion criteria and clarification behavior.
|
|
347
|
+
|
|
348
|
+
Runtime state must not be committed as config. The schema rejects known runtime
|
|
349
|
+
keys in committed config files.
|
|
350
|
+
|
|
351
|
+
## Current Boundaries
|
|
352
|
+
|
|
353
|
+
The package currently focuses on local runner workflows:
|
|
354
|
+
|
|
355
|
+
- explicit one-off runs;
|
|
356
|
+
- daemon polling;
|
|
357
|
+
- project-local config;
|
|
358
|
+
- runner-owned worktree cleanup;
|
|
359
|
+
- GitHub Issues and Pull Requests;
|
|
360
|
+
- Codex CLI as the agent backend.
|
|
361
|
+
|
|
362
|
+
Hosted infrastructure, non-GitHub issue trackers, and non-Codex agents are not
|
|
363
|
+
part of the current package, although the code keeps adapter boundaries for
|
|
364
|
+
future expansion.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "codex-orchestrator",
|
|
3
|
-
"version": "0.1.
|
|
3
|
+
"version": "0.1.25",
|
|
4
4
|
"description": "Reusable GitHub Issues runner for Codex.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"license": "MIT",
|
|
@@ -22,6 +22,7 @@
|
|
|
22
22
|
"files": [
|
|
23
23
|
"dist/src",
|
|
24
24
|
"prompts",
|
|
25
|
+
"docs/deep-dive.md",
|
|
25
26
|
"CHANGELOG.md",
|
|
26
27
|
"README.md",
|
|
27
28
|
"LICENSE"
|
|
@@ -11,6 +11,12 @@ Run code-review before completion for runtime changes and report the result.
|
|
|
11
11
|
Report validation, skipped checks, and risks.
|
|
12
12
|
For UI or visual changes, follow the orchestration prompt's visual proof
|
|
13
13
|
contract. If a runner-owned visual proof command is configured, prepare its
|
|
14
|
-
script/artifacts (prefer Playwright) but let the runner
|
|
15
|
-
|
|
16
|
-
|
|
14
|
+
script/artifacts (prefer Playwright for browser/web UI) but let the runner
|
|
15
|
+
execute it. For Android mobile app UI work, use device-backed proof instead of
|
|
16
|
+
Playwright: run `adb devices -l`, prefer a connected non-emulator device serial,
|
|
17
|
+
and run `export ANDROID_SERIAL=<serial>`. Otherwise run `emulator -list-avds`,
|
|
18
|
+
start an AVD in a separate shell with `emulator -avd <avd-name>`, and wait with
|
|
19
|
+
`adb wait-for-device`. If Test Android Apps skills are unavailable, try to enable
|
|
20
|
+
or load that plugin through the available Codex plugin/tool discovery mechanism.
|
|
21
|
+
If the plugin cannot be enabled, or no usable device or emulator is available,
|
|
22
|
+
report that as a warning/skipped check with the concrete reason and proceed.
|