orchestrator-workflow 0.33.0 → 0.35.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -7,6 +7,74 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.35.0] - 2026-09-15
11
+
12
+ - The npm tarball now ships a `LICENSE` file matching the repo root LICENSE
13
+ (MIT), asserted by the monorepo's `lint-package-licenses` CI job.
14
+
15
+ - A repo-level `lint-release-changelogs` CI job (`scripts/check-release-changelogs.mjs`
16
+ at the agent-dx root, named in `CONTRIBUTING.md`'s "Releasing okf-kit" steps)
17
+ fails a release commit whose package.json version bump has no matching top
18
+ `## [x.y.z]` CHANGELOG heading, whose fresh heading (checked against a
19
+ `--base <ref>`, the PR base sha or the previous push commit in CI) still
20
+ leaves `[Unreleased]` non-empty, or whose `docs/okf/log.md` entry carries a
21
+ `CHANGELOG.md:<n>` mention with no okf-kit anchor. Motivated by agent-tasks
22
+ task e077bcd8: a prior release commit (#266) bumped a version without
23
+ cutting its CHANGELOG unnoticed by CI, and two log.md entries had written
24
+ a live, unanchored `CHANGELOG.md:<n>` citation that silently resolved to
25
+ unrelated content after the next cut. Does not change okf-kit's own
26
+ citation rules or any historical CHANGELOG content.
27
+
28
+ - Hardened `lint-release-changelogs` (review round 2 on the same task):
29
+ rule 2 now fires only when the head version is a strictly greater
30
+ semver release than `--base` (a head older than base, or a version
31
+ neither side parses as strict semver, no longer misfires); an invalid
32
+ `--base` (an unresolvable ref, or a flag-shaped value) is now a usage
33
+ error instead of a silently skipped comparison. Added an
34
+ `empty-release-section` rule for a cut heading with no body before the
35
+ next heading, and a `checked-package-scope` rule pinning the four
36
+ packages this repo already checks so one silently losing its
37
+ CHANGELOG.md is a hard failure instead of a shrinking count; packages
38
+ skipped for carrying no CHANGELOG.md are now printed by name. Added a
39
+ `--root <dir>` option (a disposable fixture tree instead of this repo)
40
+ backing a new fixture-based test suite, and `CONTRIBUTING.md`'s step 6
41
+ now names `--base origin/master` explicitly instead of implying the
42
+ `[Unreleased]` check runs without it.
43
+
44
+ - Reduced the contracts reference's duplicate implementer prose to an explicit
45
+ route to workflow step 6 and the installed implementer role prompt. The
46
+ unchanged YAML contract remains the schema surface; conformance tests pin
47
+ the route, target workflow heading, and role-local verification/commit
48
+ obligations.
49
+
50
+ - Verification sets combine trusted preflight execution with ordered repository
51
+ extras through one named, immutable reference. Both roles replay the complete
52
+ set and preserve every child result, limitation, and non-pass outcome; repos
53
+ with `docs/okf/` always include bundle validation. The README carries a
54
+ parseable checked-in-list example; no execution engine or preflight checks
55
+ are added. Evidence: Pandora batch 51, D-024/D-025—four briefs omitted the
56
+ bundle check, and CI caught drift that required one additional repair round.
57
+ - The installed skill now has a compact orchestrator entrypoint with explicit
58
+ routed references for contracts, evidence/probes, review/recovery, and
59
+ run-state/harness behavior. Persisted probe plans are optional
60
+ runner-supported artifacts: immutable plan identity and mutant locator may
61
+ be delegated by reference, but only a checked-state result is evidence.
62
+ Missing/stale references block proof; intentional supersession is recorded,
63
+ never silently rewritten. Recovery records a cursor in existing run state
64
+ for invalid returns, incomplete probes, interrupted work, and bounded
65
+ repeated findings without changing waiver authority or requiring a CLI
66
+ schema.
67
+ - The review template now records each reviewer's returned method in the machine-readable `method-applied[<round>]` marker directly below its `review-method` declaration. The marker is read by grounding-mcp's completeness reader; the accompanying `Method:` line remains its constrained prose fallback. The detailed workflow reference requires the orchestrator to write every returned value and resupply an omission or mismatch before acceptance. See agent-grounding #235.
68
+ ## [0.34.0] - 2026-09-13
69
+
70
+ - Reviewer findings now carry `introduced_by_delta: yes | no | unknown`.
71
+ A `no` attribution requires a named base build and replay in `reproduction`;
72
+ it remains in the ordinary finding gate and Findings table, while only `yes`
73
+ and `unknown` participate in bounded halt and escalation rules. The legacy
74
+ five-cell table and placeholder row remain byte-compatible with the
75
+ completeness reader; concrete rows record attribution parenthetically in
76
+ their Description field.
77
+
10
78
  ## [0.33.0] - 2026-09-12
11
79
 
12
80
  ### Changed
package/INSTALL-AGENT.md CHANGED
@@ -33,7 +33,10 @@ which is mutable. For a stable audit, pin the URL to a commit SHA instead
33
33
  If the installer reports conflicts with locally edited files, inspect the
34
34
  concrete files and reuse any overwrite authority already granted for that
35
35
  scope. Ask before a `--force` re-run only when authority or conflict scope
36
- remains unresolved. **The operator
36
+ remains unresolved. The compact skill core and its required `references/`
37
+ files are one coherent bundle: the installer checks every destination for
38
+ conflicts before it activates a new core, and keeps the current bundle
39
+ unless an authorized `--force` replaces the affected files. **The operator
37
40
  path**: when an operator has already run `orchestrator-workflow setup`
38
41
  on this machine (an operator manifest exists at
39
42
  `<operator home>/manifest.json`, where the operator home is
@@ -168,6 +171,9 @@ steps in the repository you were asked to install into.
168
171
  on a re-run. If the command reports conflicts, inspect the concrete files,
169
172
  reuse prior overwrite authority for the same scope, and ask before
170
173
  `--force` only when authority or scope remains unresolved.
174
+ The conflict check covers the compact core and the complete required
175
+ `references/` set before replacing either, so a non-forced re-run preserves
176
+ the current coherent bundle.
171
177
 
172
178
  **Operator path**: before running `init`, check whether an operator
173
179
  manifest already exists on this machine, at
@@ -193,6 +199,11 @@ steps in the repository you were asked to install into.
193
199
  must run inline and sequentially until the automated CLI can generate
194
200
  `.codex/agents/*.toml`.
195
201
 
202
+ Before writing any skill core manually, check its destination and every
203
+ required `references/*.md` destination for local conflicts. If any conflict
204
+ exists, preserve the current core/reference bundle; write replacements only
205
+ with the same explicit overwrite authority required by the automated path.
206
+
196
207
  - `.ai/workflow/templates/00-goal.md` through `06-handoff.md` from
197
208
  `assets/templates/`, unchanged.
198
209
  - `.ai/runs/.gitkeep`, empty. The orchestrator later writes a
@@ -205,7 +216,8 @@ steps in the repository you were asked to install into.
205
216
  `<!-- orchestrator-workflow:begin -->` / `<!-- orchestrator-workflow:end -->`
206
217
  markers.
207
218
  - Claude Code: `.claude/skills/orchestrator-workflow/SKILL.md` from
208
- `assets/skill/SKILL.md`. For each role in the chosen profile (all five
219
+ `assets/skill/SKILL.md`, plus every regular Markdown file in
220
+ `assets/skill/references/` at the matching `references/` path. For each role in the chosen profile (all five
209
221
  for `full`; only `implementer` and `reviewer` for `minimal`),
210
222
  `.claude/agents/<role>.md` from
211
223
  `assets/agents/<role>.md` with `model: <operator's choice>` added as a
@@ -219,11 +231,12 @@ steps in the repository you were asked to install into.
219
231
  `disallowedTools: Edit, Write, NotebookEdit` goes on a new line
220
232
  directly after the `effort:` line. Ensure `CLAUDE.md` exists and
221
233
  contains a line `@AGENTS.md`.
222
- - Codex: `.agents/skills/orchestrator-workflow/SKILL.md`, same skill file.
234
+ - Codex: `.agents/skills/orchestrator-workflow/SKILL.md` and the same
235
+ `references/*.md` files.
223
236
  Do not hand-author `.codex/agents/*.toml`; report the native-agent
224
237
  limitation above and use the inline/sequential role fallback.
225
238
  - opencode: `.opencode/skills/orchestrator-workflow/SKILL.md` from
226
- `assets/skill/SKILL.md`, unchanged.
239
+ `assets/skill/SKILL.md`, plus matching `references/*.md`, unchanged.
227
240
  For each role in the chosen profile (same set as Claude Code above),
228
241
  `.opencode/agents/<role>.md` from `assets/agents/<role>.md`, with the
229
242
  frontmatter rewritten to this order: `description:` (unchanged), then
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Lan Nguyen Si
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/README.md CHANGED
@@ -99,6 +99,12 @@ automatic upgrade. [INSTALL-AGENT.md](INSTALL-AGENT.md) makes the write surface
99
99
  and fallback behavior auditable. The link tracks `master`; pin it to a commit
100
100
  SHA for a stable audit.
101
101
 
102
+ The compact skill entrypoint and its routed references form one installed
103
+ bundle. On a reinstall, the installer checks the core and every required
104
+ reference for local conflicts before activating a new core; it leaves the
105
+ current coherent bundle intact unless an explicitly authorized `--force` run
106
+ replaces the affected files.
107
+
102
108
  ### Manual and advanced CLI installation
103
109
 
104
110
  ```bash
@@ -150,6 +156,56 @@ harness to install it.
150
156
  npx orchestrator-workflow init --harness none --yes
151
157
  ```
152
158
 
159
+ ## Verification sets
160
+
161
+ A repository may check in `.ai/workflow/verify.json` to name the complete
162
+ verification set for an implementer or reviewer briefing. This generic worked
163
+ example uses a `preflight run <repo> --json` executor and ordered extras with
164
+ `cwd`, `argv`, and an explicit before/after-preflight phase, so an approved
165
+ build can precede a dependent test:
166
+
167
+ ```json
168
+ {
169
+ "format": "orchestrator-workflow-verification-set/v1",
170
+ "preflight": {
171
+ "kind": "preflight",
172
+ "name": "preflight",
173
+ "cwd": ".",
174
+ "argv": ["preflight", "run", ".", "--json"]
175
+ },
176
+ "extras": [
177
+ {
178
+ "kind": "command",
179
+ "name": "build",
180
+ "phase": "before_preflight",
181
+ "cwd": "packages/example",
182
+ "argv": ["npm", "run", "build"]
183
+ },
184
+ {
185
+ "kind": "command",
186
+ "name": "package-tests",
187
+ "phase": "after_preflight",
188
+ "cwd": "packages/example",
189
+ "argv": ["npm", "test"]
190
+ },
191
+ {
192
+ "kind": "bundlecheck",
193
+ "name": "knowledge-bundle",
194
+ "phase": "after_preflight",
195
+ "cwd": "packages/example",
196
+ "argv": ["npx", "okf-kit", "check", "docs/okf"]
197
+ }
198
+ ]
199
+ }
200
+ ```
201
+
202
+ The workflow does not execute or validate this file: the orchestrator first
203
+ approves the resolved effective config and scripts, then records a run-local
204
+ snapshot with the set digest, repository identity, executable identity, and
205
+ every result. Preflight JSON reports check results, not the underlying shell
206
+ commands it discovered. Repositories with `docs/okf/` include their bundle
207
+ check in every set, even when the task did not edit documentation.
208
+
153
209
  ## What gets installed
154
210
 
155
211
  ```text
@@ -167,11 +223,16 @@ it to the repository's `.gitignore`.
167
223
 
168
224
  Per selected harness:
169
225
 
226
+ Each installed skill includes the compact `SKILL.md` entrypoint and every
227
+ regular Markdown file from its adjacent `references/` directory. The entrypoint
228
+ routes run-state/harness, contracts, evidence/probes, and review/recovery work
229
+ to those files; references are part of the installed skill, not optional docs.
230
+
170
231
  | Harness | Files | Notes |
171
232
  |---|---|---|
172
- | Claude Code | `.claude/skills/orchestrator-workflow/SKILL.md`, `.claude/agents/{explorer,task-slicer,implementer,reviewer,advisor}.md`, `CLAUDE.md` | Claude Code reads `CLAUDE.md`, not `AGENTS.md`; the installer adds an additive `@AGENTS.md` import. Subagent models go into the `model:` frontmatter; the read-only explorer, reviewer, and advisor also get `disallowedTools: Edit, Write, NotebookEdit`. |
173
- | OpenAI Codex | `.agents/skills/orchestrator-workflow/SKILL.md`, `.codex/agents/{explorer,task-slicer,implementer,reviewer,advisor}.toml` | Codex reads `AGENTS.md` natively. Native custom-agent files carry the canonical role instructions plus `model` and `model_reasoning_effort`. Explorer and advisor request a read-only sandbox; reviewer inherits the caller's sandbox so it can run temporary/build checks, while its prompt prohibits source edits. |
174
- | opencode | `.opencode/skills/orchestrator-workflow/SKILL.md`, `.opencode/agents/{explorer,task-slicer,implementer,reviewer,advisor}.md` | opencode reads `AGENTS.md` natively. Subagents get `mode: subagent`; the read-only explorer, reviewer, and advisor also get `permission: edit: deny`. Model resolution is described below. |
233
+ | Claude Code | `.claude/skills/orchestrator-workflow/{SKILL.md,references/*.md}`, `.claude/agents/{explorer,task-slicer,implementer,reviewer,advisor}.md`, `CLAUDE.md` | Claude Code reads `CLAUDE.md`, not `AGENTS.md`; the installer adds an additive `@AGENTS.md` import. Subagent models go into the `model:` frontmatter; the read-only explorer, reviewer, and advisor also get `disallowedTools: Edit, Write, NotebookEdit`. |
234
+ | OpenAI Codex | `.agents/skills/orchestrator-workflow/{SKILL.md,references/*.md}`, `.codex/agents/{explorer,task-slicer,implementer,reviewer,advisor}.toml` | Codex reads `AGENTS.md` natively. Native custom-agent files carry the canonical role instructions plus `model` and `model_reasoning_effort`. Explorer and advisor request a read-only sandbox; reviewer inherits the caller's sandbox so it can run temporary/build checks, while its prompt prohibits source edits. |
235
+ | opencode | `.opencode/skills/orchestrator-workflow/{SKILL.md,references/*.md}`, `.opencode/agents/{explorer,task-slicer,implementer,reviewer,advisor}.md` | opencode reads `AGENTS.md` natively. Subagents get `mode: subagent`; the read-only explorer, reviewer, and advisor also get `permission: edit: deny`. Model resolution is described below. |
175
236
 
176
237
  **Read-only posture, honestly stated.** Claude Code disables file-mutation
177
238
  tools for explorer, reviewer, and advisor; opencode denies edits for those
@@ -35,6 +35,21 @@ Rules:
35
35
  gate's threshold and pass/fail counts, not a run-specific coverage
36
36
  percentage; cite a percentage only together with the exact commit and the
37
37
  run count, since branch coverage can vary between runs of the same commit.
38
+ - Run the complete repository-bound `verification_set` named in your briefing.
39
+ Before acquiring preflight output or running an extra, require the
40
+ orchestrator's approval of the resolved repository configuration and every
41
+ script/argument; the set is not authority to execute repository data. Use
42
+ the frozen run-local snapshot (set path/digest, repository identity/revision
43
+ and dirty state, effective config/scripts, and preflight executable
44
+ identity/definition). Report each executor, extra, and raw preflight child
45
+ by `(kind, name, occurrence)`, in order, with cwd and result artifact.
46
+ Preserve a missing-tool preflight limitation even when it has no child
47
+ result. A missing/extra/mismatched/unresolved result is a misfire; a failure
48
+ is reported honestly; `skip`, `acknowledged`, limitation, and inconclusive
49
+ are non-passes. A disabled required category is a gap. Always include the
50
+ bundle check when the repository has `docs/okf/`, even for unrelated edits.
51
+ Put every complete-set result in `tests.executed`, preserving the existing
52
+ report envelope for both v1 and original-contract runs.
38
53
  - When the task assignment names mutation probes to run, run each one and
39
54
  report it in the `mutation_probes` field of your output (mutant, file,
40
55
  anchor, before, after, verified_applied_via, result, expectation,
@@ -72,6 +87,14 @@ Rules:
72
87
  `result` alone is not: report it as such (`result` `survived` or
73
88
  `not_applicable` with the reason) and resolve it before the next
74
89
  reviewer spawn.
90
+ - A persisted probe-plan reference may stand in for a repeated inline mutant
91
+ definition when it resolves to a path plus immutable revision or hash and the
92
+ mutant locator/index. Resolve it before running; a missing, stale, or
93
+ unresolvable reference is `not_applicable` evidence that blocks the relevant
94
+ proof, not a skipped probe. Keep the legacy inline report fields unchanged:
95
+ the result still records the applied definition and restoration outcome.
96
+ Never rewrite a prior plan for new code; record intentional supersession and
97
+ rationale in run state before using a replacement.
75
98
  - When a verify runner is available, run it for the checks the acceptance
76
99
  criteria name and report its summary under `tests.executed`; when a
77
100
  mutation-probe runner is available, run the named probes through it and
@@ -52,6 +52,21 @@ Check, at minimum:
52
52
  green label; implementers cannot revise their own baseline. Compare the
53
53
  returned `criterion_evidence` references to every assigned frozen criterion;
54
54
  required empty references remain unresolved and block acceptance.
55
+ - Verification set: independently run the complete repository-bound
56
+ `verification_set` named in the briefing. Before acquisition or execution,
57
+ confirm the orchestrator approved the resolved effective configuration and
58
+ scripts; a repository set is not execution authority. Compare the frozen
59
+ snapshot's set path/digest, repository identity/revision/dirty state,
60
+ effective config/scripts, and preflight executable identity/definition.
61
+ Report every ordered `(kind, name, occurrence)` executor, extra, and raw
62
+ preflight child with cwd and result artifact. A missing tool may be a
63
+ limitation with no child, never a pass; disabled required categories are
64
+ gaps. Missing/extra/mismatched/unresolved results are misfires, while a
65
+ reported failure remains an honest failure. `skip`, `acknowledged`,
66
+ limitation, and inconclusive outcomes are non-passes. Require the bundle
67
+ check whenever the repository has `docs/okf/`, regardless of edit scope.
68
+ Put the independent complete-set outcome in `reproduction.result`, preserving
69
+ the existing report envelope for both v1 and original-contract runs.
55
70
  - Spec compliance: does the change do what the task contract asked, fully?
56
71
  - Architecture consistency: does it fit the existing structure and idioms?
57
72
  - Edge cases: empty inputs, error paths, concurrency, encoding, limits.
@@ -73,7 +88,7 @@ Check, at minimum:
73
88
  review round, classify each finding as `new` or `repeated` against the
74
89
  earlier rounds you were told about; on a first round every finding is
75
90
  `new` by definition. The orchestrator uses this to detect the
76
- review-round escalation budget's trigger.
91
+ review-round escalation budget's trigger. Delta attribution: classify every finding as `introduced_by_delta: yes | no | unknown`; set `no` only after naming the base build and replaying the same reproduction in `reproduction`, and record it in `05-review-findings.md` through the ordinary gate rather than bounded-round halt/escalation guidance (yes/unknown only).
77
92
  - GitHub Actions shell replay: for any diff that adds or changes a GitHub
78
93
  Actions `run:` step, replay it yourself under the shell the step actually
79
94
  runs: `bash --noprofile --norc -eo pipefail` when `shell: bash` is set on
@@ -143,9 +158,14 @@ Rules:
143
158
  through it instead of editing files by hand, and carry its result fields
144
159
  into your findings and `reproduction`; when a verify runner is available,
145
160
  read its summary before opening full logs.
161
+ - A reviewer briefing may identify a replayed probe through a resolved
162
+ immutable probe-plan reference (path plus revision/hash and mutant
163
+ locator/index) rather than repeat its inline definition. Verify the plan and
164
+ result bind the checked state, cwd, attempt, expectation, application, and
165
+ restoration; a plan alone, stale reference, or unresolved reference is not
166
+ evidence. Legacy inline probe reports remain valid.
146
167
 
147
168
  Return exactly this structure as your final output, nothing else:
148
-
149
169
  ```yaml
150
170
  status: reviewed
151
171
  role: reviewer
@@ -158,6 +178,7 @@ findings:
158
178
  description: ""
159
179
  suggested_fix: ""
160
180
  recurrence: new | repeated
181
+ introduced_by_delta: yes | no | unknown
161
182
  acceptance_recommendation: accept | accept_with_notes | fix_required | reject
162
183
  missing_tests:
163
184
  - ""
@@ -44,6 +44,12 @@ Rules:
44
44
  boundaries for the task — which files or areas the implementer may touch
45
45
  and must not touch — not implementation instructions. Apply Contract
46
46
  selection above to those fields for a recorded original contract.
47
+ - Include a repository-bound `verification_set` reference in every implementer
48
+ and reviewer briefing: its checked-in path, repository identity, and
49
+ run-local frozen snapshot. The orchestrator approves effective config and
50
+ scripts before any preflight acquisition or command execution; the set does
51
+ not grant that authority. Include an ordered bundle check whenever the
52
+ repository has `docs/okf/`, regardless of task scope.
47
53
  - For every identifier, config value, build context, or documented command a
48
54
  task will change, enumerate every file and doc site that references it in
49
55
  `relevant_files` or `relevant_docs`, with an annotation for a site the task
@@ -87,6 +93,8 @@ tasks:
87
93
  - ""
88
94
  dependencies:
89
95
  - ""
96
+ verification_set:
97
+ reference: ""
90
98
  risk: low | medium | high
91
99
  recommended_order:
92
100
  - T-001
@@ -128,9 +128,9 @@ trivial change.
128
128
  the exhausted tier path falls straight to the merge-hold), or an
129
129
  operator merge-hold, and adds a row (task, choice, reason) to
130
130
  `03-decisions.md`'s Review-round escalation table, then sets the
131
- `review-round-escalation` marker to the most recent choice. A counted
132
- round is a completed reviewer return recommending `fix_required` or
133
- `reject`; a misfired review is not a round. Which of the three is
131
+ `review-round-escalation` marker to the most recent choice. A negative round
132
+ has an `acceptance_recommendation` of `fix_required` or `reject`; a misfired
133
+ review is not a round. A negative round counts only with at least one introduced_by_delta yes/unknown finding; no stays ordinary gate. Which of the three is
134
134
  picked is judgment; that one is picked and recorded is not. Escalating
135
135
  never substitutes for a review round and comes in addition to the halt
136
136
  rule's split-or-redesign response, not instead of it.