@walwal-harness/cli 7.1.53 → 7.1.55

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,7 +4,7 @@ docmeta:
4
4
  title: walwal-harness CHANGELOG
5
5
  type: output
6
6
  createdAt: 2026-05-07T00:00:00Z
7
- updatedAt: 2026-06-24T00:00:00Z
7
+ updatedAt: 2026-09-01T00:00:00Z
8
8
  source:
9
9
  producer: agent
10
10
  skillId: harness-release
@@ -31,6 +31,18 @@ docmeta:
31
31
 
32
32
  ## Unreleased
33
33
 
34
+ ## 7.1.55 — Read before plan: lessons-before-planning ordering + Stop-hook gate (2026-09-01)
35
+ - **Hard Rule 20 — Lessons precede planning.** `AGENTS.md` now carries an *ordering* rule, not another content artifact: before any source edit, any measurement, and before any brief is issued, every CXX and every hired worker reads `.harness/conventions/{shared,role}.md` + `.harness/gotchas/{shared,role}.md`, follows only the topic links those files name, and writes `## Lessons Preflight` (which items apply and why). Every `ceo.md`, `{cxx}.md`, and worker report closes with a one-line `## Lessons Tally` immediately above `## Implementation Notes`; **`0 fired` is valid and must be stated, not omitted** — a tally that only ever reports hits trains agents to manufacture them. The rule also states the propagation clause: a requirement placed on a CXX that its workers must also satisfy is inserted **verbatim** into the worker brief, because *a rule stated one layer above the layer that executes it does not apply*. Origin: a 24h autonomous run in which the company kept re-learning lessons it had already written down — index compliance was already perfect (40 gotchas, 0 orphans, all linked), and the knowledge still arrived after the mistake. **Deliberately not shipped: the distilled 28-item preflight file** that run also produced — a derived corpus must be re-synced whenever any source file changes, went factually stale within an hour, and would be a second thing nobody reads before planning.
36
+ - **Hard Rule 21 — every worker spawn declares its model.** Never inherit the CLI default. The model is recorded in the brief, the Worker Evidence Manifest, and `progress.json company_state.workers[]`. A worker terminated by a usage limit is indistinguishable, from the outside, from a worker that finished, so **a silent or truncated worker is a rate limit until proven otherwise**. `scripts/harness-worker-dispatch.sh` now resolves an explicit model (`company_mode.worker_model` → `flow.team.{gen,eval}_model` → `opus`), passes `--model` on every `claude -p` spawn form, and records it in the queue, the dispatch record, and `progress.log`.
37
+ - **Hard Rule 22 / `conventions/shared.md` — section-scoped readers do not filter by content.** Any reader scanning a role document or worker report matches `^>?\s*#{1,6}` and returns every hit — blockquoted or plain, at any depth, no content filter. Measured against real CXX documents, a verdict-token filter caught 3/11 blockquoted headings and a retraction-marker filter 4/11; their union still missed 4, including the *second line of a verdict* whose first line was caught. A filter that catches half a verdict is worse than one that catches none, because it reports success — and a reader anchored on `^#` returns a claim while never reaching its in-place retraction. `scripts/harness-worker-evidence-validate.sh` heading matching is fixed accordingly.
38
+ - **Report skeletons are seeded at file creation, not assembled at the end.** Every CXX skill gained a `Worker Spawn Contract`: create the worker report **before** the worker starts with all required sections present (`## Status`, `## Task`, `## Evidence`, `## Result`, `## Lessons Tally`, and the stubbed terminal `## Implementation Notes`), and brief the worker to fill it in incrementally. `init` now installs the skeleton to `.harness/shared/templates/worker-report.md`. Evidence: same failure, opposite outcome — unseeded workers killed mid-round left stubs and halted the company twice; a seeded worker killed by the same rate limit left an intact report and cost nothing.
39
+ - **CQO/OPS: an instrument declares what it cannot see.** New `Instrument Validity` sections require a positive control that fires **in the same run** and varies the exact variable under suspicion before any negative evidence is admissible (otherwise the verdict is BLOCKED, not PASS), and require dependency-supplied filtering to be **read from source and quoted** (package, version, file, line range) rather than inferred from observed output. Origin: every traffic audit in the source repo inherited an invisible `skip: res.statusCode < 400` from `live-server@1.2.2` at `logLevel === 2` — a rule with 0 references anywhere in the project's own code, so every "no stubbed 2xx was served" claim rested on a structurally null instrument.
40
+ - **The Stop hook enforces it.** New `scripts/harness-lessons-gate.sh`, wired into `scripts/harness-stop.sh` alongside the existing worker-evidence guard: a turn cannot end while the latest active mission's role documents lack `## Lessons Preflight` or `## Lessons Tally`. Rules that live only in prose are followed when convenient; the Stop hook was the single most reliable actor in the source run, so the check goes where the enforcement already works. Scoped to `latest-active` so legacy/archive documents cannot block forever, and opt-out via `.harness/config.json` `behavior.lessons_gate:false`. `migrate` injects `behavior.lessons_gate` + `company_mode.worker_model` into existing configs.
41
+ - New gotcha entries carrying the run's lessons: written-indexed-and-still-too-late, silent-worker-is-a-rate-limit, heading-readers-that-filter-by-content, rules-stated-one-layer-above (shared); negative-evidence-without-a-positive-control, null-instrument-supplied-by-a-dependency, audit-questions-that-offer-alternatives (cqo); stub-report-from-a-worker-killed-mid-round (cto); clean-log-from-an-unproven-instrument (ops).
42
+ - **Fixed: every boolean config opt-out was inert.** `jq`'s `//` treats `false` as empty, so `.behavior.auto_chain_on_stop // true` returned `true` even when the project had explicitly set it to `false`. Setting `auto_chain_on_stop:false` (Stop-hook auto-chaining), `company_mode.write_on_signal:false` (legacy always-write hourly review), or `behavior.auto_route_ceo:false` (CEO routing guidance) silently did nothing. All four sites now test for `null` explicitly. Found while verifying the new `lessons_gate` opt-out, which had inherited the same shape.
43
+ - Also shipping in this release (already on the working tree): unattended wake/dispatch agents pass `--dangerously-skip-permissions` (`--dangerously-bypass-approvals-and-sandbox` for Codex) so an hourly tick never blocks on a permission prompt nobody is there to answer, and `init` pre-accepts the workspace-trust gate plus `enableAllProjectMcpServers` in `~/.claude.json` (atomic write, one-time backup, non-fatal on failure) — the two gates a project-level `settings.json` and `--dangerously-skip-permissions` cannot clear on their own.
44
+ - `/goal`, `/submission`, `/hot-fix` repeat the ordering rule at the Owner entrypoint. `/hot-fix` scales the read to the fix: a four-line patch pays the index files plus only the matching topic links, never the whole corpus — the ordering constraint holds at every size, the depth does not.
45
+
34
46
  ## 7.1.53 — Write-on-signal documents + `archive/` namespace reclaim (2026-06-24)
35
47
  - **`archive/` is now reserved for completed missions, never migrate/init backups.** Every backup snapshot moved out of `.harness/archive/` into a new `.harness/backups/` root: `bin/init.js` (migration snapshot dir, legacy-role migration, pre-harness doc extraction backup, AGENTS/CLAUDE doc backup), `init.sh refresh-ref`, `scripts/init-agents-md.sh`, and `scripts/init-ref-docs.sh`. The Owner expected `archive/` to hold a completed `goal + submission + hot-fix` per goal folder; the harness using it as a migrate backup dump conflicted with that meaning. Files are moved (not deleted), so recovery data is preserved. (Agreement: "archive/ 네임스페이스 탈환".)
36
48
  - `.harness/backups/` is added to the managed `.gitignore` block (`assets/templates/gitignore-append.txt`), and **`migrate` now refreshes the host `.gitignore` and best-effort untracks newly-ignored runtime paths** (previously only `init` did this), so migrating projects also stop committing backup/runtime noise.
@@ -13,6 +13,17 @@ Own design strategy for the mission.
13
13
 
14
14
  Before design work, read `.harness/conventions/shared.md`, `.harness/conventions/cdo.md`, `.harness/gotchas/shared.md`, and `.harness/gotchas/cdo.md`. Then follow only the related links in `cdo.md` files that match the mission topic. Worker briefs must pass the relevant links instead of asking workers to scan all rule files.
15
15
 
16
+ ### Lessons Before Plan
17
+
18
+ That read happens **before** the first source edit, the first measurement, and the first worker brief — not alongside them, and not after. The corpus is rarely the problem; the ordering is. Then, in `cdo.md`, write:
19
+
20
+ - `## Lessons Preflight` — which convention/gotcha items apply to this mission and why, named by id or heading. Written before any worker is dispatched. If the corpus genuinely has nothing for this topic, say so explicitly.
21
+ - `## Lessons Tally` — one line, written last, naming which of those items actually fired. **`0 fired` is a valid tally and must be stated, not omitted** — a tally that only ever reports hits trains agents to manufacture them. Place it immediately before `## Implementation Notes`.
22
+
23
+ **Propagate verbatim.** Any requirement this skill places on CDO that its workers must also satisfy — the linked corpus items, the browser-automation clause, the seeded report skeleton, the `## Lessons Tally` line, the `## Implementation Notes` block — is copied **word for word** into every worker brief. *A rule stated one layer above the layer that executes it does not apply,* and a worker cannot infer a rule it was never given.
24
+
25
+ Do not distill the corpus into a private checklist file and read that instead. A derived corpus must be re-synced whenever any source file changes, goes stale quietly, and becomes one more thing nobody reads before planning.
26
+
16
27
  ## Workflow
17
28
 
18
29
  1. Read CEO and COO mission context.
@@ -29,7 +40,7 @@ Before design work, read `.harness/conventions/shared.md`, `.harness/conventions
29
40
 
30
41
  When CEO routes this mission to you, set yourself as the live agent on entry so the dashboard shows the handoff: `bash scripts/harness-progress-set.sh . '.current_agent="cdo" | .agent_status="running"'`.
31
42
 
32
- Before launching any fresh worker session, update `.harness/progress.json` with `scripts/harness-progress-set.sh` so dashboards can show the worker as active. Record the worker name, owning CXX, report path, and `status:"running"` under `company_state.workers`, increment `company_state.active_workers`, and set `conductor.current_action` to `spawn:{worker-name}`. After the worker report is accepted, update that worker to `status:"complete"` and decrement `active_workers`. Do not leave `active_workers:0` while a worker session is running. Require every worker report to open with a `## Status` line whose body is `IN_PROGRESS` while the worker runs and `COMPLETE` once the report is final, so the dashboard shows true worker liveness instead of guessing from file timestamps.
43
+ Before launching any fresh worker session, update `.harness/progress.json` with `scripts/harness-progress-set.sh` so dashboards can show the worker as active. Record the worker name, owning CXX, report path, **the model the worker was spawned with**, and `status:"running"` under `company_state.workers`, increment `company_state.active_workers`, and set `conductor.current_action` to `spawn:{worker-name}`. After the worker report is accepted, update that worker to `status:"complete"` and decrement `active_workers`. Do not leave `active_workers:0` while a worker session is running. Require every worker report to open with a `## Status` line whose body is `IN_PROGRESS` while the worker runs and `COMPLETE` once the report is final, so the dashboard shows true worker liveness instead of guessing from file timestamps.
33
44
 
34
45
  On exit, after writing `cdo.md` and handing back to CEO, run `bash scripts/harness-progress-set.sh . '.agent_status="completed"'` so the loop advances and the dashboard reflects the finished step. Do not clear `conductor.state`; only the CEO's Company Loop Termination step ends the loop.
35
46
 
@@ -59,11 +70,23 @@ Owner is the final acceptance reviewer, not a design QA substitute. CDO must use
59
70
 
60
71
  Required output sections:
61
72
 
62
- 1. Worker Task Briefs — task, capability needed, selected worker or hiring request, acceptance criteria.
63
- 2. Worker Evidence Manifest — worker name, report path, status.
64
- 3. CDO Decision — only decisions accepted from worker evidence.
65
- 4. Preview Artifact — path `.harness/documents/{mission_name}/cdo/preview.html`, visual summary, and dashboard viewing note.
66
- 5. Next Handoff — CTO-ready design constraints, inputs, blockers.
67
- 6. Implementation Notes — in English, with `Design Decisions`, `Deviations`, `Tradeoffs`, and `Open Questions`.
73
+ 1. Lessons Preflight — convention/gotcha items that apply to this mission, why each applies, and the topic links passed into worker briefs. Written before the first worker is dispatched.
74
+ 2. Worker Task Briefs — task, capability needed, selected worker or hiring request, declared model, acceptance criteria.
75
+ 3. Worker Evidence Manifest — worker name, declared model, report path, status.
76
+ 4. CDO Decision — only decisions accepted from worker evidence.
77
+ 5. Preview Artifact — path `.harness/documents/{mission_name}/cdo/preview.html`, visual summary, and dashboard viewing note.
78
+ 6. Next Handoff — CTO-ready design constraints, inputs, blockers.
79
+ 7. Lessons Tally — one line naming which preflight items actually fired. `0 fired` is valid and must be stated.
80
+ 8. Implementation Notes — in English, with `Design Decisions`, `Deviations`, `Tradeoffs`, and `Open Questions`.
68
81
 
69
82
  Every CDO worker brief must require the worker to append the same English `## Implementation Notes` block to the bottom of `.harness/documents/{mission_name}/cdo/workers/{worker-name}.md`, covering risks, self-corrections, chosen direction, and unresolved questions. Use `None` for empty subsections.
83
+
84
+ ## Worker Spawn Contract
85
+
86
+ Two things are decided **before** the round starts, not after a worker dies.
87
+
88
+ **1. Declare the model.** Every worker spawn names its model explicitly — never inherit the CLI or session default. Record that model in the brief, in the Worker Evidence Manifest, and in `company_state.workers[]`. A worker terminated by a usage limit is indistinguishable, from the outside, from a worker that finished, so **a silent or truncated worker is a rate limit until proven otherwise**: check the limit and its reset time before re-briefing, re-hiring, or rewriting the task. Spreading a round across model families is only a decision you can make if the model was declared.
89
+
90
+ **2. Seed the report.** Create `.harness/documents/{goal-or-child-mission}/cdo/workers/{worker-name}.md` **before the worker starts**, already carrying every required section — `## Status` (`IN_PROGRESS`), `## Task`, `## Evidence`, `## Result`, `## Lessons Tally`, and the terminal `## Implementation Notes` block with all four subsections stubbed. Copy `.harness/shared/templates/worker-report.md` when it is installed; otherwise write the skeleton by hand. Brief the worker to fill it in **incrementally as the work happens**, never to assemble the report at the end.
91
+
92
+ A worker that dies mid-round — rate limit, crash, cancelled session — must leave a **valid partial report, never a stub**. Same failure, opposite outcome, one variable: unseeded workers killed mid-round left stubs and halted the company; a seeded worker killed by the same limit left its report intact and cost nothing. The variable was a decision taken before the round.
@@ -72,6 +72,16 @@ Before routing or accepting CXX work, enforce lazy loading:
72
72
  - Topic files such as `.harness/gotchas/i18n-locale-hotfix.md` remain separate. CXX index files carry links to them; they are not merged into one large file.
73
73
  - Workers receive the relevant CXX link set in their brief instead of scanning every convention/gotcha file.
74
74
 
75
+ ### Lessons Before Plan
76
+
77
+ The read is an **ordering constraint**, not a reading list. It happens before the first source edit, the first measurement, and the first brief — not alongside them, and not after. A lesson that is written, indexed, and reachable still arrives too late if it is read after the mistake.
78
+
79
+ - `ceo.md` opens with `## Lessons Preflight`: which convention/gotcha items apply to this mission and why. Write it before routing to the first CXX.
80
+ - `ceo.md` closes with a one-line `## Lessons Tally` immediately before `## Implementation Notes`: which of those items actually fired. **`0 fired` is a valid tally and must be stated, not omitted.**
81
+ - CEO requires the same two sections from every `{cxx}.md` and every worker report, and **rejects** any report that omits them.
82
+ - **Propagate verbatim.** Any requirement CEO places on a CXX that its workers must also satisfy is copied word for word into the worker brief by that CXX. *A rule stated one layer above the layer that executes it does not apply.* CEO checks the worker briefs recorded in `{cxx}.md` for this, not just the CXX document.
83
+ - **Do not commission a distilled preflight checklist.** A derived corpus must be re-synced whenever any source file changes, goes stale silently, and becomes a second thing nobody reads before planning. Fix the ordering, not the corpus.
84
+
75
85
  ## Mission Protocol
76
86
 
77
87
  1. Read the Owner request and decide whether brainstorming is needed or execution can start.
@@ -103,6 +113,9 @@ Before routing or accepting CXX work, enforce lazy loading:
103
113
  - CEO must not accept CQO PASS for a runnable product unless OPS has supplied clean verification-watch evidence or an explicit not-applicable reason. Open OPS incidents, missing runtime mapping, missing required logs, service down, or health mismatch block Owner acceptance.
104
114
  - After launch, CEO treats OPS production incidents as company events. CEO convenes CTO/CQO/OPS when user-impacting production signals appear; CTO owns recovery, CQO owns regression confirmation, and OPS owns evidence and close criteria.
105
115
  - Every CEO and CXX mission document must include an English `## Implementation Notes` section with the required subsections below. CEO must reject CXX reports that omit it.
116
+ - Every CEO and CXX mission document and every worker report must carry `## Lessons Preflight` and a one-line `## Lessons Tally` (tally immediately before `## Implementation Notes`). `0 fired` is a valid tally; an omitted tally is not. CEO must reject reports that omit either section.
117
+ - Every worker spawn declares its model explicitly — never the inherited CLI default. CEO requires the model in each CXX's Worker Task Briefs and Worker Evidence Manifest. **A silent or truncated worker is a rate limit until proven otherwise**: before treating a stalled round as a CXX failure, CEO asks for the usage-limit status and reset time, because a rate-limited worker looks exactly like a finished one from the outside.
118
+ - Every worker report file is created **before the worker starts**, already carrying its required sections (seeded from `.harness/shared/templates/worker-report.md`). A worker killed mid-round must leave a valid partial report, never a stub. CEO treats a stub report as a seeding failure by the owning CXX, not as a worker failure.
106
119
  - Owner is the final acceptance reviewer, not a tester, QA substitute, debugger, or deployment verifier. CEO must not send "done, please check" reports while core functionality, regression, account setup, browser flows, logs, or runtime health remain unverified by workers.
107
120
  - Before requesting Owner acceptance, CEO must collect and summarize CXX-backed completion evidence: CTO implementation evidence, CQO evaluator/tester evidence, and OPS verification-watch/runtime evidence when runnable environments are involved. The final Owner report may request acceptance review or business/product judgment, but must not ask the Owner to discover whether the software works.
108
121
 
@@ -126,6 +139,13 @@ Every `ceo.md` and CXX document (`coo.md`, `cdo.md`, `cto.md`, `cqo.md`, `ops.md
126
139
  - ...
127
140
  ```
128
141
 
142
+ Immediately above it, every document carries the one-line tally:
143
+
144
+ ```
145
+ ## Lessons Tally
146
+ - Fired: <items from Lessons Preflight that actually changed a decision, or `0 fired`>
147
+ ```
148
+
129
149
  Use `None` when a subsection has no entries. These notes are mandatory even for small or emergency work. They must summarize how the role interpreted the Owner request, where the role intentionally diverged from the request, what alternatives were considered, and any true external-authority blocker. Do not list routine CEO-approved operations as "needs Owner confirmation."
130
150
 
131
151
  When briefing a CXX, CEO must explicitly require the CXX to append this section to its own `{cxx}.md` and to require every worker it manages to append the same section to the bottom of that worker's report.
@@ -13,6 +13,17 @@ Own mission planning, research, references, hypotheses, and goal fit.
13
13
 
14
14
  Before planning, read `.harness/conventions/shared.md`, `.harness/conventions/coo.md`, `.harness/gotchas/shared.md`, and `.harness/gotchas/coo.md`. Then follow only the related links in `coo.md` files that match the mission topic. Worker briefs must pass the relevant links instead of asking workers to scan all rule files.
15
15
 
16
+ ### Lessons Before Plan
17
+
18
+ That read happens **before** the first source edit, the first measurement, and the first worker brief — not alongside them, and not after. The corpus is rarely the problem; the ordering is. Then, in `coo.md`, write:
19
+
20
+ - `## Lessons Preflight` — which convention/gotcha items apply to this mission and why, named by id or heading. Written before any worker is dispatched. If the corpus genuinely has nothing for this topic, say so explicitly.
21
+ - `## Lessons Tally` — one line, written last, naming which of those items actually fired. **`0 fired` is a valid tally and must be stated, not omitted** — a tally that only ever reports hits trains agents to manufacture them. Place it immediately before `## Implementation Notes`.
22
+
23
+ **Propagate verbatim.** Any requirement this skill places on COO that its workers must also satisfy — the linked corpus items, the browser-automation clause, the seeded report skeleton, the `## Lessons Tally` line, the `## Implementation Notes` block — is copied **word for word** into every worker brief. *A rule stated one layer above the layer that executes it does not apply,* and a worker cannot infer a rule it was never given.
24
+
25
+ Do not distill the corpus into a private checklist file and read that instead. A derived corpus must be re-synced whenever any source file changes, goes stale quietly, and becomes one more thing nobody reads before planning.
26
+
16
27
  ## MCP Capability Scan
17
28
 
18
29
  During planning, COO must determine whether the active runtime exposes MCP servers, MCP tools, connector tools, or tool-discovery tools that can materially improve the mission. This is a planning input, not an implementation shortcut.
@@ -39,7 +50,7 @@ During planning, COO must determine whether the active runtime exposes MCP serve
39
50
 
40
51
  When CEO routes this mission to you, set yourself as the live agent on entry so the dashboard shows the handoff: `bash scripts/harness-progress-set.sh . '.current_agent="coo" | .agent_status="running"'`.
41
52
 
42
- Before launching any fresh worker session, update `.harness/progress.json` with `scripts/harness-progress-set.sh` so dashboards can show the worker as active. Record the worker name, owning CXX, report path, and `status:"running"` under `company_state.workers`, increment `company_state.active_workers`, and set `conductor.current_action` to `spawn:{worker-name}`. After the worker report is accepted, update that worker to `status:"complete"` and decrement `active_workers`. Do not leave `active_workers:0` while a worker session is running. Require every worker report to open with a `## Status` line whose body is `IN_PROGRESS` while the worker runs and `COMPLETE` once the report is final, so the dashboard shows true worker liveness instead of guessing from file timestamps.
53
+ Before launching any fresh worker session, update `.harness/progress.json` with `scripts/harness-progress-set.sh` so dashboards can show the worker as active. Record the worker name, owning CXX, report path, **the model the worker was spawned with**, and `status:"running"` under `company_state.workers`, increment `company_state.active_workers`, and set `conductor.current_action` to `spawn:{worker-name}`. After the worker report is accepted, update that worker to `status:"complete"` and decrement `active_workers`. Do not leave `active_workers:0` while a worker session is running. Require every worker report to open with a `## Status` line whose body is `IN_PROGRESS` while the worker runs and `COMPLETE` once the report is final, so the dashboard shows true worker liveness instead of guessing from file timestamps.
43
54
 
44
55
  On exit, after writing `coo.md` and handing back to CEO, run `bash scripts/harness-progress-set.sh . '.agent_status="completed"'` so the loop advances and the dashboard reflects the finished step. Do not clear `conductor.state`; only the CEO's Company Loop Termination step ends the loop.
45
56
 
@@ -72,11 +83,23 @@ Return planning decisions, evidence, rejected options, mission fit, worker names
72
83
 
73
84
  Required output sections:
74
85
 
75
- 1. Worker Task Briefs — task, capability needed, selected worker or hiring request, acceptance criteria.
76
- 2. Worker Evidence Manifest — worker name, report path, status.
77
- 3. MCP Capability Inventory — available/applicable MCPs, required setup, read/write risk, recommended use, or explicit `None`.
78
- 4. COO Decision — only decisions accepted from worker evidence.
79
- 5. Next Handoff — next CXX, inputs, blockers.
80
- 6. Implementation Notes — in English, with `Design Decisions`, `Deviations`, `Tradeoffs`, and `Open Questions`.
86
+ 1. Lessons Preflight — convention/gotcha items that apply to this mission, why each applies, and the topic links passed into worker briefs. Written before the first worker is dispatched.
87
+ 2. Worker Task Briefs — task, capability needed, selected worker or hiring request, declared model, acceptance criteria.
88
+ 3. Worker Evidence Manifest — worker name, declared model, report path, status.
89
+ 4. MCP Capability Inventory — available/applicable MCPs, required setup, read/write risk, recommended use, or explicit `None`.
90
+ 5. COO Decision — only decisions accepted from worker evidence.
91
+ 6. Next Handoff — next CXX, inputs, blockers.
92
+ 7. Lessons Tally — one line naming which preflight items actually fired. `0 fired` is valid and must be stated.
93
+ 8. Implementation Notes — in English, with `Design Decisions`, `Deviations`, `Tradeoffs`, and `Open Questions`.
81
94
 
82
95
  Every COO worker brief must require the worker to append the same English `## Implementation Notes` block to the bottom of `.harness/documents/{mission_name}/coo/workers/{worker-name}.md`, covering risks, self-corrections, chosen direction, and unresolved questions. Use `None` for empty subsections.
96
+
97
+ ## Worker Spawn Contract
98
+
99
+ Two things are decided **before** the round starts, not after a worker dies.
100
+
101
+ **1. Declare the model.** Every worker spawn names its model explicitly — never inherit the CLI or session default. Record that model in the brief, in the Worker Evidence Manifest, and in `company_state.workers[]`. A worker terminated by a usage limit is indistinguishable, from the outside, from a worker that finished, so **a silent or truncated worker is a rate limit until proven otherwise**: check the limit and its reset time before re-briefing, re-hiring, or rewriting the task. Spreading a round across model families is only a decision you can make if the model was declared.
102
+
103
+ **2. Seed the report.** Create `.harness/documents/{goal-or-child-mission}/coo/workers/{worker-name}.md` **before the worker starts**, already carrying every required section — `## Status` (`IN_PROGRESS`), `## Task`, `## Evidence`, `## Result`, `## Lessons Tally`, and the terminal `## Implementation Notes` block with all four subsections stubbed. Copy `.harness/shared/templates/worker-report.md` when it is installed; otherwise write the skeleton by hand. Brief the worker to fill it in **incrementally as the work happens**, never to assemble the report at the end.
104
+
105
+ A worker that dies mid-round — rate limit, crash, cancelled session — must leave a **valid partial report, never a stub**. Same failure, opposite outcome, one variable: unseeded workers killed mid-round left stubs and halted the company; a seeded worker killed by the same limit left its report intact and cost nothing. The variable was a decision taken before the round.
@@ -13,6 +13,17 @@ Own quality, recurrence prevention, and archive eligibility.
13
13
 
14
14
  Before quality work, read `.harness/conventions/shared.md`, `.harness/conventions/cqo.md`, `.harness/gotchas/shared.md`, and `.harness/gotchas/cqo.md`. Then follow only the related links in `cqo.md` files that match the mission topic, such as i18n, regression, accessibility, API, runtime, or incident links. Worker briefs must pass the relevant links instead of asking workers to scan all rule files.
15
15
 
16
+ ### Lessons Before Plan
17
+
18
+ That read happens **before** the first source edit, the first measurement, and the first worker brief — not alongside them, and not after. The corpus is rarely the problem; the ordering is. Then, in `cqo.md`, write:
19
+
20
+ - `## Lessons Preflight` — which convention/gotcha items apply to this mission and why, named by id or heading. Written before any worker is dispatched. If the corpus genuinely has nothing for this topic, say so explicitly.
21
+ - `## Lessons Tally` — one line, written last, naming which of those items actually fired. **`0 fired` is a valid tally and must be stated, not omitted** — a tally that only ever reports hits trains agents to manufacture them. Place it immediately before `## Implementation Notes`.
22
+
23
+ **Propagate verbatim.** Any requirement this skill places on CQO that its workers must also satisfy — the linked corpus items, the browser-automation clause, the seeded report skeleton, the `## Lessons Tally` line, the `## Implementation Notes` block — is copied **word for word** into every worker brief. *A rule stated one layer above the layer that executes it does not apply,* and a worker cannot infer a rule it was never given.
24
+
25
+ Do not distill the corpus into a private checklist file and read that instead. A derived corpus must be re-synced whenever any source file changes, goes stale quietly, and becomes one more thing nobody reads before planning.
26
+
16
27
  ## Workflow
17
28
 
18
29
  1. Read CEO and CTO mission context.
@@ -28,7 +39,7 @@ Before quality work, read `.harness/conventions/shared.md`, `.harness/convention
28
39
 
29
40
  When CEO routes this mission to you, set yourself as the live agent on entry so the dashboard shows the handoff: `bash scripts/harness-progress-set.sh . '.current_agent="cqo" | .agent_status="running"'`.
30
41
 
31
- Before launching any fresh worker session, update `.harness/progress.json` with `scripts/harness-progress-set.sh` so dashboards can show the worker as active. Record the worker name, owning CXX, report path, and `status:"running"` under `company_state.workers`, increment `company_state.active_workers`, and set `conductor.current_action` to `spawn:{worker-name}`. After the worker report is accepted, update that worker to `status:"complete"` and decrement `active_workers`. Do not leave `active_workers:0` while a worker session is running. Require every worker report to open with a `## Status` line whose body is `IN_PROGRESS` while the worker runs and `COMPLETE` once the report is final, so the dashboard shows true worker liveness instead of guessing from file timestamps.
42
+ Before launching any fresh worker session, update `.harness/progress.json` with `scripts/harness-progress-set.sh` so dashboards can show the worker as active. Record the worker name, owning CXX, report path, **the model the worker was spawned with**, and `status:"running"` under `company_state.workers`, increment `company_state.active_workers`, and set `conductor.current_action` to `spawn:{worker-name}`. After the worker report is accepted, update that worker to `status:"complete"` and decrement `active_workers`. Do not leave `active_workers:0` while a worker session is running. Require every worker report to open with a `## Status` line whose body is `IN_PROGRESS` while the worker runs and `COMPLETE` once the report is final, so the dashboard shows true worker liveness instead of guessing from file timestamps.
32
43
 
33
44
  On exit, after writing `cqo.md` and handing back to CEO, run `bash scripts/harness-progress-set.sh . '.agent_status="completed"'` so the loop advances and the dashboard reflects the finished step. Do not clear `conductor.state`; only the CEO's Company Loop Termination step ends the loop.
34
45
 
@@ -36,6 +47,33 @@ On exit, after writing `cqo.md` and handing back to CEO, run `bash scripts/harne
36
47
 
37
48
  When the active goal is operating (perpetual, `mission-state.json` lifecycle `operating`), CEO periodically orders a 현황 보고. In it, confirm — with evaluator-worker evidence — whether the live system still passes the quality/regression bar toward the goal (for CQO: are regression/e2e/perf/security gates still green on the running system?). If you discover a regression, quality drift, incident, or verification gap, do not silently sit on it: raise it as an agenda item so CEO can adjudicate and route the next cycle: `bash scripts/harness-agenda.sh . <goal-rel> raise cqo <kind> "<title>" "<evidence-path>"` (kinds: loss, drift, incident, opportunity, risk, verification-gap). When CEO routes a decided agenda item to you, run the evaluator/tester workers, issue a verdict, and report so CEO can close the item.
38
49
 
50
+ ## Test Coverage Scope And Full Gate
51
+
52
+ CQO must verify the mission in two layers:
53
+
54
+ 1. Changed-scope verification: evaluator/tester workers inspect and test only the files, modules, APIs, flows, and adjacent dependencies identified in CTO's handoff.
55
+ 2. Final full-suite gate: CQO runs the project's full test/coverage command once, near the end, through an evaluator/tester worker and normal project tooling.
56
+
57
+ CQO must not ask evaluator workers to manually perform full-project test coverage analysis by LLM inspection. The full gate must use fast executable tooling such as `npm test`, `npm run test:coverage`, `pnpm test`, `pytest`, `go test ./...`, CI-equivalent scripts, or the repository's documented command. If no full-suite command exists, CQO records that as a verification gap instead of inventing a manual full-coverage review.
58
+
59
+ If changed-scope tests or changed-scope coverage fail, CQO returns FAIL or BLOCKED for CTO correction. If the final full-suite gate fails or reports coverage gaps outside the changed scope, CQO must classify it as one of:
60
+
61
+ - Side-effect suspected: changed work appears to have broken unrelated behavior.
62
+ - Out-of-scope pre-existing gap: failure or coverage deficit is unrelated to the mission changes.
63
+ - Inconclusive: insufficient evidence to distinguish side effect from pre-existing state.
64
+
65
+ CQO must report any out-of-scope full-suite failure or coverage deficit to CEO and CTO with command output, affected paths, and the classification above. CQO must not expand the mission into broad unrelated test-writing work unless CEO explicitly routes that as a new task.
66
+
67
+ ## Instrument Validity
68
+
69
+ Evidence about what did **not** happen is worth exactly as much as the instrument that looked for it.
70
+
71
+ - **Negative evidence is inadmissible without a positive control that fires in the same run**, and the control must vary the exact variable under suspicion. "No error was logged", "no stubbed 2xx was served", "no leak was detected" are claims about the instrument until a control proves the instrument can see the thing at all. Require the control in the evaluator brief, not after the fact.
72
+ - Report the control next to the result: what was injected, that it was observed, and the negative result from the same run. A verdict resting on unproven negative evidence is **BLOCKED**, not PASS.
73
+ - **Where an instrument is supplied by a dependency rather than written in-repo, its filtering behaviour is read from source and quoted** — package, version, file, line range — not inferred from observed output. A filter that lives upstream is invisible to every in-repo search, so its absence from the project's own code is not evidence of its absence.
74
+ - When an instrument turns out to have been structurally null, the claims it produced are identifiable **by their shape** — every claim of that form, not just the one that happened to be noticed. Re-open them as a class and say so in Recurrence Notes.
75
+ - An audit question that offers alternatives asserts that the alternatives are exhaustive. "Is it A or B?" cannot return "neither, it is upstream". When an audit stalls, re-ask the question without the menu.
76
+
39
77
  ## Hard Rules
40
78
 
41
79
  CQO must not directly execute QA, visual review, security review, performance testing, or regression checks. CQO may only define gates, select evaluators, review evidence, decide archive eligibility, and document worker names and report paths.
@@ -59,12 +97,15 @@ If OPS reports an incident during verification:
59
97
 
60
98
  Required output sections in `cqo.md`:
61
99
 
62
- 1. Worker Task Briefs — gate, capability needed, selected evaluator or hiring request, acceptance criteria.
63
- 2. Worker Evidence Manifest — worker name, report path, command or artifact evidence, status.
64
- 3. OPS Watch Evidence — ops report path, monitored runtime mapping, incidents/warnings, and whether runtime evidence permits PASS.
65
- 4. CQO Verdict — PASS, FAIL, or BLOCKED based only on worker evidence plus required OPS watch evidence. Must reference Worker Evidence Manifest entries.
66
- 5. Recurrence Notes — accepted gotchas, conventions, memories, or none.
67
- 6. Implementation Notes — in English, with `Design Decisions`, `Deviations`, `Tradeoffs`, and `Open Questions`.
100
+ 1. Lessons Preflight — convention/gotcha items that apply to this mission, why each applies, and the topic links passed into evaluator briefs. Written before the first evaluator is dispatched.
101
+ 2. Worker Task Briefs — gate, capability needed, selected evaluator or hiring request, declared model, acceptance criteria.
102
+ 3. Worker Evidence Manifest — worker name, declared model, report path, command or artifact evidence, status.
103
+ 4. Instrument Validity — for every negative claim: the instrument, its log level and filter (quoted from source when the instrument comes from a dependency), and the positive control that fired in the same run. Negative evidence with no control is BLOCKED, not PASS.
104
+ 5. OPS Watch Evidence — ops report path, monitored runtime mapping, incidents/warnings, and whether runtime evidence permits PASS.
105
+ 6. CQO Verdict — PASS, FAIL, or BLOCKED based only on worker evidence plus required OPS watch evidence. Must reference Worker Evidence Manifest entries.
106
+ 7. Recurrence Notes — accepted gotchas, conventions, memories, or none.
107
+ 8. Lessons Tally — one line naming which preflight items actually fired. `0 fired` is valid and must be stated.
108
+ 9. Implementation Notes — in English, with `Design Decisions`, `Deviations`, `Tradeoffs`, and `Open Questions`.
68
109
 
69
110
  ## Worker Report Note Requirement
70
111
 
@@ -87,3 +128,13 @@ Every CQO evaluator/tester brief must require the worker to append this English
87
128
  ```
88
129
 
89
130
  The worker notes must cover risks, self-corrections, and chosen direction. Use `None` when a subsection has no entries. CQO must not accept evaluator output that omits this block.
131
+
132
+ ## Worker Spawn Contract
133
+
134
+ Two things are decided **before** the round starts, not after a worker dies.
135
+
136
+ **1. Declare the model.** Every worker spawn names its model explicitly — never inherit the CLI or session default. Record that model in the brief, in the Worker Evidence Manifest, and in `company_state.workers[]`. A worker terminated by a usage limit is indistinguishable, from the outside, from a worker that finished, so **a silent or truncated worker is a rate limit until proven otherwise**: check the limit and its reset time before re-briefing, re-hiring, or rewriting the task. Spreading a round across model families is only a decision you can make if the model was declared.
137
+
138
+ **2. Seed the report.** Create `.harness/documents/{goal-or-child-mission}/cqo/workers/{worker-name}.md` **before the worker starts**, already carrying every required section — `## Status` (`IN_PROGRESS`), `## Task`, `## Evidence`, `## Result`, `## Lessons Tally`, and the terminal `## Implementation Notes` block with all four subsections stubbed. Copy `.harness/shared/templates/worker-report.md` when it is installed; otherwise write the skeleton by hand. Brief the worker to fill it in **incrementally as the work happens**, never to assemble the report at the end.
139
+
140
+ A worker that dies mid-round — rate limit, crash, cancelled session — must leave a **valid partial report, never a stub**. Same failure, opposite outcome, one variable: unseeded workers killed mid-round left stubs and halted the company; a seeded worker killed by the same limit left its report intact and cost nothing. The variable was a decision taken before the round.
@@ -13,6 +13,17 @@ Own engineering execution for the mission.
13
13
 
14
14
  Before engineering work, read `.harness/conventions/shared.md`, `.harness/conventions/cto.md`, `.harness/gotchas/shared.md`, and `.harness/gotchas/cto.md`. Then follow only the related links in `cto.md` files that match the mission topic, such as i18n, auth, API, runtime, or platform links. Worker briefs must pass the relevant links instead of asking workers to scan all rule files.
15
15
 
16
+ ### Lessons Before Plan
17
+
18
+ That read happens **before** the first source edit, the first measurement, and the first worker brief — not alongside them, and not after. The corpus is rarely the problem; the ordering is. Then, in `cto.md`, write:
19
+
20
+ - `## Lessons Preflight` — which convention/gotcha items apply to this mission and why, named by id or heading. Written before any worker is dispatched. If the corpus genuinely has nothing for this topic, say so explicitly.
21
+ - `## Lessons Tally` — one line, written last, naming which of those items actually fired. **`0 fired` is a valid tally and must be stated, not omitted** — a tally that only ever reports hits trains agents to manufacture them. Place it immediately before `## Implementation Notes`.
22
+
23
+ **Propagate verbatim.** Any requirement this skill places on CTO that its workers must also satisfy — the linked corpus items, the browser-automation clause, the seeded report skeleton, the `## Lessons Tally` line, the `## Implementation Notes` block — is copied **word for word** into every worker brief. *A rule stated one layer above the layer that executes it does not apply,* and a worker cannot infer a rule it was never given.
24
+
25
+ Do not distill the corpus into a private checklist file and read that instead. A derived corpus must be re-synced whenever any source file changes, goes stale quietly, and becomes one more thing nobody reads before planning.
26
+
16
27
  ## Workflow
17
28
 
18
29
  1. Read CEO, COO, and CDO mission documents.
@@ -30,7 +41,7 @@ Before engineering work, read `.harness/conventions/shared.md`, `.harness/conven
30
41
 
31
42
  When CEO routes this mission to you, set yourself as the live agent on entry so the dashboard shows the handoff: `bash scripts/harness-progress-set.sh . '.current_agent="cto" | .agent_status="running"'`.
32
43
 
33
- Before launching any fresh worker session, update `.harness/progress.json` with `scripts/harness-progress-set.sh` so dashboards can show the worker as active. Record the worker name, owning CXX, report path, and `status:"running"` under `company_state.workers`, increment `company_state.active_workers`, and set `conductor.current_action` to `spawn:{worker-name}`. After the worker report is accepted, update that worker to `status:"complete"` and decrement `active_workers`. Do not leave `active_workers:0` while a worker session is running. Require every worker report to open with a `## Status` line whose body is `IN_PROGRESS` while the worker runs and `COMPLETE` once the report is final, so the dashboard shows true worker liveness instead of guessing from file timestamps.
44
+ Before launching any fresh worker session, update `.harness/progress.json` with `scripts/harness-progress-set.sh` so dashboards can show the worker as active. Record the worker name, owning CXX, report path, **the model the worker was spawned with**, and `status:"running"` under `company_state.workers`, increment `company_state.active_workers`, and set `conductor.current_action` to `spawn:{worker-name}`. After the worker report is accepted, update that worker to `status:"complete"` and decrement `active_workers`. Do not leave `active_workers:0` while a worker session is running. Require every worker report to open with a `## Status` line whose body is `IN_PROGRESS` while the worker runs and `COMPLETE` once the report is final, so the dashboard shows true worker liveness instead of guessing from file timestamps.
34
45
 
35
46
  On exit, after writing `cto.md` and handing back to CEO, run `bash scripts/harness-progress-set.sh . '.agent_status="completed"'` so the loop advances and the dashboard reflects the finished step. Do not clear `conductor.state`; only the CEO's Company Loop Termination step ends the loop.
36
47
 
@@ -46,6 +57,22 @@ Every CTO worker brief that may use Playwright, browser automation, browser-base
46
57
 
47
58
  CTO must not accept worker plans or reports that omit this requirement when browser automation is in scope.
48
59
 
60
+ ## Test Coverage Scope
61
+
62
+ CTO must optimize engineering verification around the work actually changed in the mission. CTO worker briefs must require targeted tests, coverage checks, and regression commands for the changed files, modules, APIs, flows, and directly affected dependencies only.
63
+
64
+ CTO must not require workers to manually reason through full-project test coverage, inspect unrelated coverage gaps, or chase 100% coverage outside the modified scope. Full-project test execution belongs to CQO's final gate and must be run by project test tooling, not by LLM inspection.
65
+
66
+ When handing off to CQO, CTO must include:
67
+
68
+ - Changed files and affected modules.
69
+ - Targeted test commands already run by workers.
70
+ - Coverage evidence for the changed scope.
71
+ - Known risk areas and directly adjacent dependencies.
72
+ - Suggested full-suite command if the project exposes one, such as `npm test`, `npm run test:coverage`, `pnpm test`, `pytest`, `go test ./...`, or the repo's equivalent.
73
+
74
+ If targeted verification fails inside the changed scope, CTO blocks the handoff until workers fix or explicitly document the blocker. If unrelated tests or coverage gaps are noticed outside the changed scope, CTO records them as possible pre-existing risk or side-effect signal and routes them through CEO/CQO instead of expanding the implementation mission by default.
75
+
49
76
  ## Hard Rules
50
77
 
51
78
  CTO must not directly write code, create build scripts, choose detailed implementation content, run technical QA as the evaluator, or produce final implementation artifacts. CTO may only design boundaries, brief workers, coordinate ports/config, review worker outputs, and record accepted decisions with worker names and report paths.
@@ -58,12 +85,14 @@ Every worker dispatched by CTO must be listed in the Worker Evidence Manifest se
58
85
 
59
86
  Required output sections in `cto.md`:
60
87
 
61
- 1. Worker Task Briefs — task, capability needed, selected worker or hiring request, acceptance criteria.
62
- 2. Port And Runtime Contract — `.env` and `.harness/config.json` values that workers must update or use.
63
- 3. Worker Evidence Manifest — worker name, report path, changed files or artifact paths, status.
64
- 4. CTO Decision — only decisions accepted from worker evidence.
65
- 5. CQO Handoff — validation scope, commands, risk areas, blockers.
66
- 6. Implementation Notes — in English, with `Design Decisions`, `Deviations`, `Tradeoffs`, and `Open Questions`.
88
+ 1. Lessons Preflight — convention/gotcha items that apply to this mission, why each applies, and the topic links passed into worker briefs. Written before the first worker is dispatched.
89
+ 2. Worker Task Briefs — task, capability needed, selected worker or hiring request, declared model, acceptance criteria.
90
+ 3. Port And Runtime Contract — `.env` and `.harness/config.json` values that workers must update or use.
91
+ 4. Worker Evidence Manifest — worker name, declared model, report path, changed files or artifact paths, status.
92
+ 5. CTO Decision — only decisions accepted from worker evidence.
93
+ 6. CQO Handoff — validation scope, commands, risk areas, blockers.
94
+ 7. Lessons Tally — one line naming which preflight items actually fired. `0 fired` is valid and must be stated.
95
+ 8. Implementation Notes — in English, with `Design Decisions`, `Deviations`, `Tradeoffs`, and `Open Questions`.
67
96
 
68
97
  ## Worker Report Note Requirement
69
98
 
@@ -86,3 +115,13 @@ Every CTO worker brief must require the worker to append this English block to t
86
115
  ```
87
116
 
88
117
  The worker notes must cover risks, self-corrections, and chosen direction. Use `None` when a subsection has no entries. CTO must not accept worker output that omits this block.
118
+
119
+ ## Worker Spawn Contract
120
+
121
+ Two things are decided **before** the round starts, not after a worker dies.
122
+
123
+ **1. Declare the model.** Every worker spawn names its model explicitly — never inherit the CLI or session default. Record that model in the brief, in the Worker Evidence Manifest, and in `company_state.workers[]`. A worker terminated by a usage limit is indistinguishable, from the outside, from a worker that finished, so **a silent or truncated worker is a rate limit until proven otherwise**: check the limit and its reset time before re-briefing, re-hiring, or rewriting the task. Spreading a round across model families is only a decision you can make if the model was declared.
124
+
125
+ **2. Seed the report.** Create `.harness/documents/{goal-or-child-mission}/cto/workers/{worker-name}.md` **before the worker starts**, already carrying every required section — `## Status` (`IN_PROGRESS`), `## Task`, `## Evidence`, `## Result`, `## Lessons Tally`, and the terminal `## Implementation Notes` block with all four subsections stubbed. Copy `.harness/shared/templates/worker-report.md` when it is installed; otherwise write the skeleton by hand. Brief the worker to fill it in **incrementally as the work happens**, never to assemble the report at the end.
126
+
127
+ A worker that dies mid-round — rate limit, crash, cancelled session — must leave a **valid partial report, never a stub**. Same failure, opposite outcome, one variable: unseeded workers killed mid-round left stubs and halted the company; a seeded worker killed by the same limit left its report intact and cost nothing. The variable was a decision taken before the round.
@@ -16,6 +16,7 @@ Hire workers from `.harness/shared/HR-Resource/`.
16
16
  - mission name
17
17
  - blocking status
18
18
  - owning CXX (`cto`, `cqo`, `coo`, `cdo`, or `ops`)
19
+ - declared model for the spawn — explicit, never the inherited CLI default
19
20
 
20
21
  ## Workflow
21
22
 
@@ -26,12 +27,24 @@ Hire workers from `.harness/shared/HR-Resource/`.
26
27
  5. Update `.harness/shared/hr-roster.json` without deleting existing hired entries. Record `owner` as the owning CXX, `skillPath` as `.harness/shared/HR-Resource/{name}/SKILL.md`, and `skillPaths.claude` / `skillPaths.codex` as tool-specific hierarchical installed paths.
27
28
  6. The owning CXX must write worker reports under `.harness/documents/{mission}/{owning-cxx}/workers/{name}.md`. Do not write flat `.harness/documents/{mission}/workers/{name}.md` except when migrating legacy missions.
28
29
  7. Ask the `harness-resource-manager` skill to update trigger wording.
29
- 8. Return worker name, owner, source skill path, installed paths, mission report path, invocation wording, related convention/gotcha links supplied by the owning CXX, and the mandatory report appendix below.
30
+ 8. Return worker name, owner, declared model, source skill path, installed paths, mission report path, invocation wording, related convention/gotcha links supplied by the owning CXX, and the mandatory report appendix below.
30
31
 
31
32
  ## Worker Rule Links
32
33
 
33
34
  Every hired worker receives the owning CXX's relevant convention/gotcha links in the worker brief. Workers read those linked topic files only when they match the assigned task.
34
35
 
36
+ Requirements the owning CXX must satisfy and that its worker must also satisfy are copied into the brief **verbatim** — the browser-automation clause, the report skeleton, the `## Lessons Tally` line, the `## Implementation Notes` block. A rule stated one layer above the layer that executes it does not apply. Reject a hire request whose brief omits them.
37
+
38
+ ## Declared Model
39
+
40
+ A hire request with no model is incomplete. Record the declared model in `hr-roster.json` alongside `owner` and the skill paths, and return it to the owning CXX for the Worker Evidence Manifest.
41
+
42
+ A worker terminated by a usage limit is indistinguishable, from the outside, from a worker that finished. When a CXX comes back asking to re-hire a worker whose round went silent, ask for the usage-limit status first: **a silent loop is a rate limit until proven otherwise**, and re-hiring on the same exhausted model family repeats the failure.
43
+
44
+ ## Seeded Report
45
+
46
+ The owning CXX creates `.harness/documents/{mission}/{owning-cxx}/workers/{name}.md` **before the worker starts**, seeded from `.harness/shared/templates/worker-report.md` with every required section already present, and briefs the worker to fill it in incrementally. A worker killed mid-round must leave a valid partial report, never a stub. Do not report a hire as complete while the report file has not been seeded.
47
+
35
48
  ## Mandatory Worker Report Appendix
36
49
 
37
50
  Every hired worker must append this English section to the bottom of its existing report:
@@ -22,6 +22,17 @@ OPS owns three environment classes:
22
22
  OPS is not the implementation owner. CTO/DevOps workers start or change systems; OPS observes whether the declared build/service environments are healthy and raises evidence-backed events.
23
23
  OPS must not directly perform DevOps implementation, service fixes, config rewrites, deployment changes, or recovery work. OPS may only monitor, classify, brief hired Ops/DevOps workers, review their reports, and escalate evidence-backed events.
24
24
 
25
+ ### Lessons Before Plan
26
+
27
+ That read happens **before** the first source edit, the first measurement, and the first worker brief — not alongside them, and not after. The corpus is rarely the problem; the ordering is. Then, in `ops.md`, write:
28
+
29
+ - `## Lessons Preflight` — which convention/gotcha items apply to this mission and why, named by id or heading. Written before any worker is dispatched. If the corpus genuinely has nothing for this topic, say so explicitly.
30
+ - `## Lessons Tally` — one line, written last, naming which of those items actually fired. **`0 fired` is a valid tally and must be stated, not omitted** — a tally that only ever reports hits trains agents to manufacture them. Place it immediately before `## Implementation Notes`.
31
+
32
+ **Propagate verbatim.** Any requirement this skill places on OPS that its workers must also satisfy — the linked corpus items, the browser-automation clause, the seeded report skeleton, the `## Lessons Tally` line, the `## Implementation Notes` block — is copied **word for word** into every worker brief. *A rule stated one layer above the layer that executes it does not apply,* and a worker cannot infer a rule it was never given.
33
+
34
+ Do not distill the corpus into a private checklist file and read that instead. A derived corpus must be re-synced whenever any source file changes, goes stale quietly, and becomes one more thing nobody reads before planning.
35
+
25
36
  ## CEO-Approved Operations
26
37
 
27
38
  OPS must not ask the Owner to approve routine monitoring operations. OPS proposes a default to CEO, and CEO decides.
@@ -68,6 +79,14 @@ After service launch, OPS continues the same monitoring duty against `runtime.pr
68
79
  - Repeated incidents must trigger recovery coordination through CEO -> CTO/CQO/OPS. OPS supplies evidence and recovery criteria; CTO owns fixes; CQO owns regression confirmation.
69
80
  - OPS may classify resolved events as close candidates only after the monitored endpoint is healthy and logs no longer show the triggering error pattern.
70
81
 
82
+ ## Instrument Validity
83
+
84
+ OPS supplies most of the harness's negative evidence — "no crash", "no error in the log", "health stayed green" — so OPS owns proving the instrument could have seen the failure.
85
+
86
+ - Every claim of the form "nothing bad happened" ships with a **positive control that fired in the same run** and varied the exact variable under suspicion. Without one, report the observation as unverified, not clean.
87
+ - **Read dependency-supplied filtering from source and quote it** — package, version, file, line range — instead of inferring it from what appeared in the log. Log middleware, dev servers, proxies, and test runners routinely drop successful or sub-threshold requests at a log level nobody chose deliberately, and that rule appears nowhere in the project's own code.
88
+ - Record the instrument in Environment Evidence: what tool observed the runtime, at what log level, with what filter, and what the control was. An unrecorded instrument makes every negative result from that run unusable.
89
+
71
90
  ## Port Policy
72
91
 
73
92
  - CEO/CTO/OPS must choose an available `{xx}000` base port before CXX services are allocated, unless the Owner already specified one.
@@ -94,7 +113,7 @@ After service launch, OPS continues the same monitoring duty against `runtime.pr
94
113
 
95
114
  When CEO routes this mission to you, set yourself as the live agent on entry so the dashboard shows the handoff: `bash scripts/harness-progress-set.sh . '.current_agent="ops" | .agent_status="running"'`.
96
115
 
97
- Before launching any fresh worker session, update `.harness/progress.json` with `scripts/harness-progress-set.sh` so dashboards can show the worker as active. Record the worker name, owning CXX, report path, and `status:"running"` under `company_state.workers`, increment `company_state.active_workers`, and set `conductor.current_action` to `spawn:{worker-name}`. After the worker report is accepted, update that worker to `status:"complete"` and decrement `active_workers`. Do not leave `active_workers:0` while a worker session is running. Require every worker report to open with a `## Status` line whose body is `IN_PROGRESS` while the worker runs and `COMPLETE` once the report is final, so the dashboard shows true worker liveness instead of guessing from file timestamps.
116
+ Before launching any fresh worker session, update `.harness/progress.json` with `scripts/harness-progress-set.sh` so dashboards can show the worker as active. Record the worker name, owning CXX, report path, **the model the worker was spawned with**, and `status:"running"` under `company_state.workers`, increment `company_state.active_workers`, and set `conductor.current_action` to `spawn:{worker-name}`. After the worker report is accepted, update that worker to `status:"complete"` and decrement `active_workers`. Do not leave `active_workers:0` while a worker session is running. Require every worker report to open with a `## Status` line whose body is `IN_PROGRESS` while the worker runs and `COMPLETE` once the report is final, so the dashboard shows true worker liveness instead of guessing from file timestamps.
98
117
 
99
118
  On exit, after writing `ops.md` and handing back to CEO, run `bash scripts/harness-progress-set.sh . '.agent_status="completed"'` so the loop advances and the dashboard reflects the finished step. Do not clear `conductor.state`; only the CEO's Company Loop Termination step ends the loop.
100
119
 
@@ -112,12 +131,24 @@ OPS must not accept worker plans or reports that omit this requirement when brow
112
131
 
113
132
  ## Required Output Sections
114
133
 
115
- 1. Worker Task Briefs — monitoring/recovery task, capability needed, selected worker or hiring request, acceptance criteria.
116
- 2. Environment Evidence — config path, command/service checked, observed status.
117
- 3. Worker Evidence Manifest — worker name, report path, status for delegated monitoring or recovery tasks.
118
- 4. OPS Event Decision — good-case silence, warning, incident, or emergency escalation.
119
- 5. CQO Verification Watch — whether CQO testing was monitored, runtime mapping used, open incidents, and PASS/BLOCKED implication.
120
- 6. Post-Launch Watch — production services monitored, open incidents, recovery status, or not applicable.
121
- 7. Implementation Notes — in English, with `Design Decisions`, `Deviations`, `Tradeoffs`, and `Open Questions`.
134
+ 1. Lessons Preflight — convention/gotcha items that apply to this mission, why each applies, and the topic links passed into worker briefs. Written before the first worker is dispatched.
135
+ 2. Worker Task Briefs — monitoring/recovery task, capability needed, selected worker or hiring request, declared model, acceptance criteria.
136
+ 3. Environment Evidence — config path, command/service checked, observed status, **and the instrument**: what tool observed it, at what log level, with what filter, and the positive control that fired in the same run.
137
+ 4. Worker Evidence Manifest — worker name, declared model, report path, status for delegated monitoring or recovery tasks.
138
+ 5. OPS Event Decision — good-case silence, warning, incident, or emergency escalation.
139
+ 6. CQO Verification Watch — whether CQO testing was monitored, runtime mapping used, open incidents, and PASS/BLOCKED implication.
140
+ 7. Post-Launch Watch — production services monitored, open incidents, recovery status, or not applicable.
141
+ 8. Lessons Tally — one line naming which preflight items actually fired. `0 fired` is valid and must be stated.
142
+ 9. Implementation Notes — in English, with `Design Decisions`, `Deviations`, `Tradeoffs`, and `Open Questions`.
122
143
 
123
144
  Every OPS worker brief must require the worker to append the same English `## Implementation Notes` block to the bottom of `.harness/documents/{mission_name}/ops/workers/{worker-name}.md`, covering risks, self-corrections, chosen direction, and unresolved questions. Use `None` for empty subsections.
145
+
146
+ ## Worker Spawn Contract
147
+
148
+ Two things are decided **before** the round starts, not after a worker dies.
149
+
150
+ **1. Declare the model.** Every worker spawn names its model explicitly — never inherit the CLI or session default. Record that model in the brief, in the Worker Evidence Manifest, and in `company_state.workers[]`. A worker terminated by a usage limit is indistinguishable, from the outside, from a worker that finished, so **a silent or truncated worker is a rate limit until proven otherwise**: check the limit and its reset time before re-briefing, re-hiring, or rewriting the task. Spreading a round across model families is only a decision you can make if the model was declared.
151
+
152
+ **2. Seed the report.** Create `.harness/documents/{goal-or-child-mission}/ops/workers/{worker-name}.md` **before the worker starts**, already carrying every required section — `## Status` (`IN_PROGRESS`), `## Task`, `## Evidence`, `## Result`, `## Lessons Tally`, and the terminal `## Implementation Notes` block with all four subsections stubbed. Copy `.harness/shared/templates/worker-report.md` when it is installed; otherwise write the skeleton by hand. Brief the worker to fill it in **incrementally as the work happens**, never to assemble the report at the end.
153
+
154
+ A worker that dies mid-round — rate limit, crash, cancelled session — must leave a **valid partial report, never a stub**. Same failure, opposite outcome, one variable: unseeded workers killed mid-round left stubs and halted the company; a seeded worker killed by the same limit left its report intact and cost nothing. The variable was a decision taken before the round.
@@ -176,7 +176,7 @@ All Playwright usage by CEO, CXX, and hired workers must run with a visible real
176
176
  8. **No CXX self-execution** — CXX agents coordinate and manage only. A CXX that produces deliverables without matching worker records has violated its scope. CEO must reject such reports.
177
177
  9. **No verdict without worker evidence** — CQO cannot issue ACCEPTED/REJECTED without a Worker Evidence Manifest referencing at least one evaluator worker. Self-inspection by CQO is not valid evidence.
178
178
  10. **Hierarchical worker ownership** — Hired workers are installed under `.claude/skills/{owning-cxx}/{worker}/` and `.codex/skills/{owning-cxx}/{worker}/`; mission worker reports live under `.harness/documents/{goal-or-child-mission}/{owning-cxx}/workers/`. Flat `{mission}/workers/` reports are legacy and signal an ownership violation unless explicitly migrated.
179
- 11. **Lazy convention/gotcha loading** — CXX roles read `.harness/conventions/shared.md`, `.harness/conventions/{cxx}.md`, `.harness/gotchas/shared.md`, and `.harness/gotchas/{cxx}.md`, then follow only the related topic links listed in those CXX files. Workers read only the related links supplied by their owning CXX.
179
+ 11. **Lazy convention/gotcha loading** — *What* to read. *When* to read it, and what to write about it, is Rule 20. CXX roles read `.harness/conventions/shared.md`, `.harness/conventions/{cxx}.md`, `.harness/gotchas/shared.md`, and `.harness/gotchas/{cxx}.md`, then follow only the related topic links listed in those CXX files. Workers read only the related links supplied by their owning CXX.
180
180
  12. **Implementation Notes required** — `ceo.md`, every `{cxx}.md`, and every worker report must end with one English `## Implementation Notes` section containing `Design Decisions`, `Deviations`, `Tradeoffs`, and `Open Questions`. Use `None` for empty subsections. Do not create a separate sidecar notes file; the notes belong at the bottom of the same role or worker report that produced the decision/evidence.
181
181
  13. **Structured runtime state** — If the harness must parse it, record it in JSON or JSONL. CXX todo queues, event history, heartbeat timestamps, preemption/resume state, and completion evidence belong in `.harness/todos/*.json*` or `.harness/events.jsonl`, not in free-form Markdown tables. Markdown remains for instructions, conventions, gotchas, skills, and human-readable mission narrative.
182
182
  14. **Mission lifecycle is explicit** — Every goal, submission, and hot-fix directory must contain `mission-state.json` with `lifecycle` and `active`. Only one child mission under a goal may be active. Starting a newer submission/hot-fix closes, cancels, or supersedes the previous active child unless CEO records a deliberate TODO/resume plan.
@@ -185,6 +185,9 @@ All Playwright usage by CEO, CXX, and hired workers must run with a visible real
185
185
  17. **Production incidents are company events** — After launch, user-impacting OPS signals route through CEO to CTO/CQO/OPS. CTO owns recovery, CQO owns regression confirmation, and OPS owns evidence plus close criteria.
186
186
  18. **Playwright is visible by default** — All Playwright/browser automation must use headed/visible mode (`headless: false` or equivalent). Headless Playwright requires explicit Owner approval recorded in the mission.
187
187
  19. **The loop ends only via a runtime transition** — There are exactly two legitimate stop conditions: COMPLETE and BLOCKED-on-external-authority. The autonomous Stop loop and the dashboard read `progress.json` runtime state, not mission documents. After the final Owner report and terminal `mission-state.json`, CEO's last action MUST be `scripts/harness-company-complete.sh` (complete) or `scripts/harness-company-block.sh "<missing authority>"` (blocked). Never end a turn in the `running` state with no queued action, and never fire a terminal transition while real CXX/worker/verification work remains. Between these two endpoints the company keeps progressing autonomously — the Owner is not the pump.
188
+ 20. **Lessons precede planning** — Before any source edit, any measurement, and before any brief is issued, every CXX and every hired worker reads `.harness/conventions/shared.md`, `.harness/conventions/{role}.md`, `.harness/gotchas/shared.md`, and `.harness/gotchas/{role}.md`, follows only the topic links those files name, and writes a `## Lessons Preflight` section stating which items apply and why. Any requirement placed on a CXX that its workers must also satisfy is inserted **verbatim** into the worker brief — *a rule stated one layer above the layer that executes it does not apply.* Every `ceo.md`, every `{cxx}.md`, and every worker report carries a one-line `## Lessons Tally` naming which of those items actually fired, placed immediately before `## Implementation Notes`. **Zero fired is a valid tally and must be stated, not omitted** — a tally that only ever reports hits trains agents to manufacture them. Do not distill the corpus into a second checklist file: a derived corpus must be re-synced whenever any source file changes, and it becomes one more thing nobody reads before planning.
189
+ 21. **Every worker spawn declares its model** — Never inherit the CLI default. The spawning CXX names the model in the brief, in the Worker Evidence Manifest, and in `progress.json` `company_state.workers[]`. A worker terminated by a usage limit is indistinguishable, from the outside, from a worker that finished — so **a silent loop is a rate limit until proven otherwise.** Check the limit and its reset time before re-briefing, re-hiring, or rewriting the task.
190
+ 22. **Readers do not filter by content** — Any reader that scans a role document, worker report, or mission record for sections matches `^>?\s*#{1,6}` and returns every hit: blockquoted or plain, at any depth, with no content filter of any kind. Do not select headings by what they appear to say. A document is free to put verdicts, retractions, standing rules, and continuation lines in a heading, so a reader anchored on `^#` can return superseded content as current — worse than returning nothing. **The reader that decides what is important before reading is the failure.**
188
191
 
189
192
  ---
190
193
 
@@ -168,7 +168,9 @@
168
168
  "behavior": {
169
169
  "comment": "하네스 동작 플래그. UserPromptSubmit 훅이 이 값을 읽어 v7 CEO/CXX 라우팅 안내를 결정한다.",
170
170
  "auto_route_ceo": true,
171
- "auto_route_ceo_description": "true 이면 /goal 또는 /hot-fix 이후 Owner 입력을 v7 CEO/CXX mission flow 기준으로 안내한다. 사용자가 'harness skip' 등을 말하면 단일 메시지 한정으로 건너뛴다."
171
+ "auto_route_ceo_description": "true 이면 /goal 또는 /hot-fix 이후 Owner 입력을 v7 CEO/CXX mission flow 기준으로 안내한다. 사용자가 'harness skip' 등을 말하면 단일 메시지 한정으로 건너뛴다.",
172
+ "lessons_gate": true,
173
+ "lessons_gate_description": "true 이면 Stop 훅(scripts/harness-lessons-gate.sh)이 활성 미션의 role 문서에 '## Lessons Preflight'(적용 conventions/gotchas 항목과 이유)와 '## Lessons Tally'(실제 발동 항목, '0 fired' 도 유효)가 없으면 턴 종료를 막는다. AGENTS.md Hard Rule 20 — 교훈은 계획보다 먼저 읽는다. 아직 채택하지 않은 프로젝트는 false 로 opt-out."
172
174
  },
173
175
  "token_limit": {
174
176
  "comment": "모델 토큰 제한으로 작업이 중단됐을 때의 저비용 재개 정책. 별도 probe 호출 없이 시간 기반으로만 재개 알림을 계산한다.",
@@ -184,6 +186,8 @@
184
186
  "worker_concurrency": 3,
185
187
  "worker_spawn": "claude",
186
188
  "worker_spawn_description": "claude 이면 idle worker 배정 시 `claude -p` 백그라운드 프로세스를 띄운다. record 로 두면 prompt/log만 생성한다.",
189
+ "worker_model": "opus",
190
+ "worker_model_description": "worker spawn 이 명시적으로 선언하는 모델. AGENTS.md Hard Rule 21 — CLI 기본 모델 상속 금지. 사용량 한도로 종료된 worker 는 밖에서 보면 정상 완료와 구별되지 않으므로, 조용한 루프는 반증 전까지 rate limit 으로 취급한다. 비우면 CLI 기본값을 쓰지만 그 선택 자체를 기록해야 한다.",
187
191
  "hourly_wake_executor": "claude",
188
192
  "hourly_wake_executor_description": "claude 이면 `claude -p`, codex 이면 `codex exec` 로 매시간 autonomous tick 을 실행한다.",
189
193
  "hourly_wake_model": "",