@walwal-harness/cli 7.1.56 → 7.1.58

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -31,6 +31,33 @@ docmeta:
31
31
 
32
32
  ## Unreleased
33
33
 
34
+ ## 7.1.58 — Mission tiers and evidence-safe completion (2026-09-28)
35
+
36
+ - Added before/after source fingerprints for verification evidence reuse, conditional on successful baselines and unchanged commands, runtime, criteria, and environment inputs.
37
+ - Added verification artifact hygiene and a secret scanner that accepts environment variable names, emits no matched values, and distinguishes leaks from incomplete scans.
38
+
39
+ - Added S/M/L routing: S CTO implements and CQO runs tests directly in a separate session; M keeps workers with compact role reports; L retains the full procedure. COO/CDO deliverables remain worker-backed at every tier.
40
+ - **CQO is a pass gate, never parallel with CTO.** CQO runs no test until CTO writes `## CQO Handoff` and exits; the handoff freezes the tree. A tree change during verification is `BLOCKED` (routed to CEO), not a silent rerun. FAIL→fix iterations rerun only failed and changed-scope tests; the full suite runs once on the final handed-off tree. This removes the loop where every CTO edit invalidated the fingerprint and restarted the full suite. The M-tier "CQO drafts gates alongside CTO" wording is removed everywhere.
41
+ - **Report only when summoned.** A CXX that CEO does not summon writes no document. OPS is summoned only when verification exercises a long-lived runtime (dev server, Docker, preview, cloud, device); self-exiting tests and builds get one `OPS N/A: <reason>` line in cqo.md. A summoned COO/CDO/OPS with no work files `## Not Applicable` plus one reason — no worker, Lessons, or Notes; the evidence and lessons gates accept it (never for CTO/CQO). At S/M, OPS observes directly without a worker; L OPS stays worker-backed.
42
+ - S/M role documents can use compact Lessons and Implementation Notes. Worker briefs reference seeded reports instead of copying their templates; hot-fix lessons are signal-based at S/M.
43
+ - Review hardening: decorated verdict labels and whitespace before the colon are invalid final candidates; nested re-evaluation and parenthesized CQO Verdict sections are included. Explicit legacy completion checks the selected mission’s lessons rather than latest-active.
44
+ - Completion accepts an explicit mission path and checks evidence before state changes. Only the last verdict candidate counts; malformed final verdicts cannot resurrect an earlier PASS. Cancelled/superseded/closed missions end without acceptance and retain their lifecycle.
45
+ - Effective tiers use numeric history ranks, including both endpoints. S upgrade exemptions require an unchanged Direct Work SHA and declared post-upgrade work. Legacy completion and the mission_tiers=false opt-out remain supported.
46
+ - Validation: 99 transition/lessons cases plus SHA/Stop/backstop checks pass; dashboard unit/company-flow tests (25) and installed sandbox tests (9) pass. Syntax checks, install-contract checks, and npm pack dry-run pass. `bash tests/mission-tiers.sh` covers verdicts, unchanged state on refusal, non-acceptance endings, OPS evidence, upgrades, legacy behavior and Stop integration. The sandbox migration test now compares the installed bundle version with package.json instead of stale 7.1.48.
47
+ - Performance comparison: the design targets 2–3 S sessions (plus OPS when applicable), versus the previous estimated 7–9. These are routing estimates, not measured results. A paired live LLM `/hot-fix` benchmark (same function bug + test, calls/time/document count and command/exit-code evidence) was not run in this implementation session; no wall-clock speedup is claimed.
48
+
49
+ ## 7.1.57 — Reading is not reaching: corpus reachability, spec pins, the document is the record (2026-09-01)
50
+ Applies the field-trial revision of the change proposal. Its P1/P4/P5/P6/P9 shipped in 7.1.55–56 and the trial marks them **field-tested** — briefed verbatim into two agents that were not told they were observed, scored on instruments fixed in advance: correct ordering, 4 real citations, 0 fabrications, and a published not-read list. **Three proposals are new, and one claim was corrected.**
51
+
52
+ - **Reachability (new, the trial's own discovery).** The trial's agent followed the reading rule exactly and still missed the entry written for its exact bug — because that entry named `cto` in its own front matter and was linked from `cqo`'s index alone. Lazy loading tells a reader to consult `{shared, own-role}`; that is a **promise about reachability**, and where it is not kept the rule does not narrow the search, **it hides the entry**. Measured: 69 items, **10 unreachable role-routings, 7 invisible to a role the entry itself named** — every one written by a single role and filed only under that role's index. New `scripts/harness-corpus-reachability.sh` audits and (`--fix`) repairs this; `harness-gotcha-register.sh` gains `--roles` and auto-links at registration, so an entry cannot be filed under its author alone. **P1 makes agents read the index; this is what makes the index worth reading.**
53
+ - **Spec pins (new).** `scripts/harness-spec-pin.sh` records the version **and content hash** of every external spec a mission builds against in `{mission}/spec-pins.json`, and re-verifies before completion and archive. This was the trial's actual defect: a spec moved `v0.7 → v0.9`, changing a response contract, while the category built against `v0.7` sat marked **complete** — two revisions had landed silently. The symptom was not an error but an absence: a lookup key stopped matching and three overlays were dropped as `null`. **A category is complete against a spec version, never in the abstract.**
54
+ - **The document is the record (new).** A conclusion a session holds but has not written into its role document and the runtime state file is not held by the company; reconcile before reporting, strike and correct in place, never delete, and check role documents against peer documents. `harness-worker-evidence-validate.sh` now catches the measured case mechanically — a worker whose report says `COMPLETE` while `progress.json` still lists it running. Cost of the unreconciled version: an orchestration loop that went on trying to spawn a finished step **70 times**.
55
+ - **Instrument validity gains a statistics clause.** A summary statistic is published **with its `n`**, and a spiky series is characterised by **percentiles, never min–max** — a range is the two least representative points in the set and reads as a finding.
56
+ - **Zero new numbered rules.** The proposal's own closing question was whether nine more entries on a nineteen-item list is P1's failure mode wearing the shape of a fix. §8 stays at 22: reachability folds into Rule 11 (which becomes *lazy loading is a promise about reachability*), spec pins into Rule 4 (*nothing is complete in the abstract*), and the record clause into Rule 12 (now *the document is the record*). The enforcement lives in scripts, not in more prose.
57
+ - Gate ordering fixed while wiring this: the spec-pin check initially sat after the runtime transition, where a refusal would have refused nothing. All three completion gates now run **before** `progress.json` is touched, verified by asserting the runtime stays `running` on refusal.
58
+ - The audit found two unreachable routings in this package's own bundled corpus on first run, and read its own syntax documentation as data until the declaration was anchored to line start. Both fixed; the bundled corpus now audits clean at install.
59
+ - On P9, the trial **corrected** the earlier claim rather than confirming it: the same verbatim rule was obeyed by the second agent, so the honest figure is one of two — not "prose does not work" but "prose works at a rate you cannot predict per agent." Still the argument for a hook, since a gate has no compliance rate, but the earlier framing overstated it.
60
+
34
61
  ## 7.1.56 — Close the lessons-gate bypass at the terminal transition (2026-09-01)
35
62
  - **Fix: the Hard Rule 20 gate could be bypassed by firing the completion transition.** Found by an A/B sandbox test of 7.1.54 vs 7.1.55 running identical fixtures. 7.1.55 put the gate in `harness-stop.sh` only, but `harness-company-complete.sh` sets `conductor.state=completed`, and `harness-stop.sh` short-circuits on that state at the top of the file — so a mission could complete having never recorded what it read, simply by running the transition. That is precisely the failure the gate was written to prevent. `harness-company-complete.sh` now runs `harness-lessons-gate.sh` itself and refuses the transition (non-zero exit, runtime left in `running`, refusal logged to `progress.log`) when the active mission's role documents lack `## Lessons Preflight` / `## Lessons Tally`.
36
63
  - Deadlock-safe by construction, verified against every legitimate stop path: the gate is scoped to the latest **active** mission, so the Stop-hook auto-complete backstop — which fires only when no mission is active — always passes; and `harness-company-block.sh` (external-authority BLOCKED) is deliberately **not** gated, since a mission blocked on a missing credential must still be able to stop.
@@ -81,6 +81,12 @@ Required output sections:
81
81
 
82
82
  Every CDO worker brief must require the worker to append the same English `## Implementation Notes` block to the bottom of `.harness/documents/{mission_name}/cdo/workers/{worker-name}.md`, covering risks, self-corrections, chosen direction, and unresolved questions. Use `None` for empty subsections.
83
83
 
84
+ ## The Document Is The Record
85
+
86
+ A conclusion you hold but have not written into `cdo.md` **is not held by the company.** Before reporting any state change — to CEO, to a peer CXX, to the Owner — reconcile it in your own document *and* in `progress.json`. Strike and correct in place; never delete the superseded line, because a reader arriving later needs to see that it was superseded rather than never written.
87
+
88
+ Check your document against your peers' documents, not only against itself. The cheap version of this failure is a deliverable table that contradicts three messages you already sent. The expensive version was measured: a completed step reported and accepted, never written to the state file, and an orchestration loop that went on trying to spawn it **70 times**.
89
+
84
90
  ## Worker Spawn Contract
85
91
 
86
92
  Two things are decided **before** the round starts, not after a worker dies.
@@ -36,8 +36,7 @@ If an operational choice is reversible and uses existing project-local credentia
36
36
 
37
37
  There are exactly two legitimate ways to end the company loop, and each requires an explicit **runtime transition**. Writing `ceo.md` and `mission-state.json` is not enough: the autonomous Stop loop and the dashboard read `progress.json` runtime state (`conductor.state`, `agent_status`, `current_agent`), not your documents. If you finish a report but never fire the transition, the loop stays `running` and the harness keeps prompting you to continue — this is the "is it done or not?" ambiguity. Avoid it by always ending in one of these two states:
38
38
 
39
- 1. **COMPLETE** — the mission is genuinely finished: the final Owner report is in `ceo.md`, all required CXX/worker/OPS evidence is collected, and `mission-state.json` is terminal (`complete`/`closed`/`cancelled`/`superseded`) with `active:false`. As the **literal final action of the turn**, run:
40
- `bash scripts/harness-company-complete.sh . <reason>`
39
+ 1. **COMPLETE / termination without acceptance** — For acceptance, leave the mission active and run `bash scripts/harness-company-complete.sh . <reason> <mission-rel>` as the final action; the script checks evidence before writing `complete` and `active:false`. On refusal, keep working. For termination without acceptance, first write lifecycle `cancelled`, `superseded`, or `closed` and `active:false`, then call the same explicit transition. `closed` means ended without acceptance; disclose “미수락 종료” in the Owner report and never archive without PASS. For an external-authority block, record `blocked`/`active:false` and run `bash scripts/harness-company-block.sh . "<exact missing authority>"`. `<mission-rel>` is the path relative to `.harness/documents/`.
41
40
  2. **BLOCKED on external authority** — the next action genuinely needs authority the harness cannot infer or obtain (new credentials/secrets, payment approval, legal/business acceptance, unavailable production access, destructive data action, or a direct conflict with the Owner's stated direction). Record the exact missing authority and the internally recommended default in `ceo.md`, set `mission-state.json` lifecycle `blocked` with `active:false`, then run:
42
41
  `bash scripts/harness-company-block.sh . "<exact missing authority>"`
43
42
 
@@ -52,7 +51,7 @@ First, **classify the goal**:
52
51
  - **Finite** goal — "build X", "add Y", "fix Z": has a definite done state. Use the Company Loop Termination above.
53
52
  - **Operating (perpetual)** goal — "run/operate/monitor/keep growing X", "지속/영구 운영", "make money continuously", anything that should *never* stop (e.g. "build a trading bot and keep it profitable forever"): it must run as a standing company that cycles indefinitely.
54
53
 
55
- For an operating goal, set `mission-state.json` to `{"lifecycle":"operating","active":true}` (it stays active forever) and **never** call `harness-company-complete.sh`. The only ways an operating goal ends are: the Owner explicitly orders it stopped (then run `harness-company-complete.sh`), or a true external-authority block (then `harness-company-block.sh`).
54
+ For an operating goal, set `mission-state.json` to `{"lifecycle":"operating","active":true,"tier":"L"}` (it stays active forever) and **never** call `harness-company-complete.sh`. The only ways an operating goal ends are: the Owner explicitly orders it stopped (then run `harness-company-complete.sh`), or a true external-authority block (then `harness-company-block.sh`).
56
55
 
57
56
  Run it as an **agenda-driven standing executive loop**. The agenda is the shared meetup file every CXX co-writes at `.harness/documents/{goal}/agenda.json`, managed with `scripts/harness-agenda.sh`. Each operating tick:
58
57
 
@@ -74,14 +73,40 @@ Before routing or accepting CXX work, enforce lazy loading:
74
73
 
75
74
  ### Lessons Before Plan
76
75
 
76
+ At effective tier S/M, role documents may replace Lessons Preflight + Lessons Tally with `## Lessons` containing `Preflight: <applicable items and why>` (before work) and `Fired: <items or 0 fired>` (at completion). `## Implementation Notes` may contain concise bullets. At L, keep the full role format below. All worker reports retain the full seeded format at every tier.
77
+
77
78
  The read is an **ordering constraint**, not a reading list. It happens before the first source edit, the first measurement, and the first brief — not alongside them, and not after. A lesson that is written, indexed, and reachable still arrives too late if it is read after the mistake.
78
79
 
79
80
  - `ceo.md` opens with `## Lessons Preflight`: which convention/gotcha items apply to this mission and why. Write it before routing to the first CXX.
80
81
  - `ceo.md` closes with a one-line `## Lessons Tally` immediately before `## Implementation Notes`: which of those items actually fired. **`0 fired` is a valid tally and must be stated, not omitted.**
81
- - CEO requires the same two sections from every `{cxx}.md` and every worker report, and **rejects** any report that omits them.
82
- - **Propagate verbatim.** Any requirement CEO places on a CXX that its workers must also satisfy is copied word for word into the worker brief by that CXX. *A rule stated one layer above the layer that executes it does not apply.* CEO checks the worker briefs recorded in `{cxx}.md` for this, not just the CXX document.
82
+ - At L, CEO requires the same two sections from every `{cxx}.md` and every worker report, and **rejects** any report that omits them.
83
+ - **Worker brief:** name the seeded report path and instruct the worker to fill its existing sections incrementally. Do not copy the report skeleton, Tally, or Notes block into the brief. Continue to pass relevant corpus links and copy behavioral requirements absent from the seed (including the browser-automation clause) verbatim.
83
84
  - **Do not commission a distilled preflight checklist.** A derived corpus must be re-synced whenever any source file changes, goes stale silently, and becomes a second thing nobody reads before planning. Fix the ordering, not the corpus.
84
85
 
86
+ ## Mission Tier
87
+
88
+ CEO records `tier: "S"|"M"|"L"` in every new mission-state.json at intake and one line of rationale in ceo.md, without asking Owner.
89
+
90
+ | Tier | Criteria and routing |
91
+ |---|---|
92
+ | S | At most 3 expected changed files OR about 150 changed lines; no new dependency or external spec; no auth, payment, security, data migration, infrastructure, or new port; existing test commands suffice. CTO implements directly; CQO verifies in a separate session. No hiring for CTO/CQO. |
93
+ | M | Remaining finite work: implementation worker(s), one evaluator worker. CQO starts only after CTO's `## CQO Handoff` (serial gate, never parallel). |
94
+ | L | Any new service/port, external spec integration, operating goal, production deployment, security, payment, or data change. Existing worker process. |
95
+
96
+ L risk criteria take precedence over size. Invoke only relevant CXX; S normally needs CTO and CQO.
97
+
98
+ **Report only when summoned.** A CXX CEO does not summon writes no document. Summon COO/CDO only for planning/UX decisions, and OPS only when verification exercises a long-lived runtime (dev server, Docker, preview, cloud, device); for self-exiting tests and builds CQO records `OPS N/A: <reason>`. A summoned COO/CDO/OPS that finds no work files `{cxx}.md` with only `## Not Applicable` and one reason line — no worker, Lessons, or Notes. COO/CDO deliverables stay worker-backed at every tier; at S/M OPS observes directly without a worker.
99
+
100
+ Effective tier is the highest rank among current `tier` and BOTH `from` and `to` in every `tier_history` entry: S=0, M=1, L=2; unknown or missing values rank 2. Never compare tier strings lexically. `behavior.mission_tiers=false` forces L (read null as true, preserve false). Missing tier means legacy: L procedure, without the new tiered completion checks.
101
+
102
+ Escalate on exceeding scope, touching a risk area, or two CQO FAILs. Set the higher tier and append `{from,to,at,reason}` to `tier_history`; downgrades never reduce effective tier. Preserve completed work at its original tier; subsequent work and final verification follow the higher tier.
103
+
104
+ For S→M/L, before changing the tier, run `bash scripts/harness-worker-evidence-validate.sh . direct-work-sha mission:<mission-rel>` and record its output as `direct_work_sha256` in that history entry. Freeze the exact `## Direct Work` body. CTO adds `## Post-Upgrade Work`: first nonblank line `none — <reason>` if only re-verification remains, otherwise assign added implementation to CTO worker(s). Missing hash removes the exemption; changed Direct Work is rejected. CQO always uses an evaluator worker after upgrade. M→L only requires expanding role documents to L format; worker reports already use the full format.
105
+
106
+ At effective tier S/M, role documents may replace Lessons Preflight + Lessons Tally with `## Lessons` containing `Preflight: <applicable items and why>` (before work) and `Fired: <items or 0 fired>` (at completion). `## Implementation Notes` may contain concise bullets. At L, keep the full role format below. All worker reports retain the full seeded format at every tier.
107
+
108
+ Across tiers: keep implementation and verification separate, use executed commands with exit codes and output excerpts, run changed-scope tests and the available full suite, never use Owner as tester, end via runtime transition only, keep Playwright headed (`headless: false`, default `slowMo: 120`), and preserve OPS watch evidence when a long-lived runtime is tested.
109
+
85
110
  ## Mission Protocol
86
111
 
87
112
  1. Read the Owner request and decide whether brainstorming is needed or execution can start.
@@ -98,10 +123,10 @@ The read is an **ordering constraint**, not a reading list. It happens before th
98
123
  ## Hard Rules
99
124
 
100
125
  - Do not let a CXX or specialist task run as an unnamed default AI engine.
101
- - CXX agents do not execute specialist work directly. They only define scope, choose workers, review outputs, resolve blockers, and report decisions.
102
- - Every mission must use hired specialist workers for research, planning, design production, implementation, QA, ops checks, or any other domain deliverable. Small scope is not an exemption.
126
+ - At M/L (and COO/CDO deliverables at every tier), CXX agents do not execute specialist work directly. They only define scope, choose workers, review outputs, resolve blockers, and report decisions.
127
+ - Effective tier determines procedural depth. Only S CTO/CQO may execute directly; the Mission Tier invariants always apply.
103
128
  - If a suitable hired worker is absent, invoke the `harness-hiring` skill before the CXX proceeds with that deliverable.
104
- - CEO must reject CXX reports that contain completed specialist deliverables without matching worker records under `.harness/documents/{goal-or-child-mission}/{owning-cxx}/workers/`.
129
+ - Except S CTO/CQO and frozen pre-upgrade CTO work, CEO must reject CXX reports that contain completed specialist deliverables without matching worker records under `.harness/documents/{goal-or-child-mission}/{owning-cxx}/workers/`.
105
130
  - Every CXX starts from fresh context and records decisions in `.harness/documents/{goal-or-child-mission}/{cxx}.md`.
106
131
  - Hiring or resource-manager output is never a stopping point. After missing workers are registered, immediately continue routing to the responsible CXX fresh sessions and require those CXX agents to brief/run the hired workers. Do not end the turn with only a hiring summary while the Owner goal remains unfinished.
107
132
  - Preserve DDD boundaries: domain decisions, application wiring, infrastructure, and quality policy are separate responsibilities.
@@ -112,8 +137,11 @@ The read is an **ordering constraint**, not a reading list. It happens before th
112
137
  - For runnable verification, collect or require CTO to record the test runtime mapping before CQO starts evaluator work: command/service name, cwd, host, port, health path if any, log path if any, and owner. OPS must watch that runtime during CQO Playwright/E2E/API/visual/performance/regression checks.
113
138
  - CEO must not accept CQO PASS for a runnable product unless OPS has supplied clean verification-watch evidence or an explicit not-applicable reason. Open OPS incidents, missing runtime mapping, missing required logs, service down, or health mismatch block Owner acceptance.
114
139
  - After launch, CEO treats OPS production incidents as company events. CEO convenes CTO/CQO/OPS when user-impacting production signals appear; CTO owns recovery, CQO owns regression confirmation, and OPS owns evidence and close criteria.
115
- - Every CEO and CXX mission document must include an English `## Implementation Notes` section with the required subsections below. CEO must reject CXX reports that omit it.
116
- - Every CEO and CXX mission document and every worker report must carry `## Lessons Preflight` and a one-line `## Lessons Tally` (tally immediately before `## Implementation Notes`). `0 fired` is a valid tally; an omitted tally is not. CEO must reject reports that omit either section.
140
+ - At L, every CEO and CXX mission document must include an English `## Implementation Notes` section with the required subsections below. CEO must reject CXX reports that omit it.
141
+ - A conclusion a CXX holds but has not written into its `{cxx}.md` **and** into `progress.json` is not held by the company. Before accepting any reported state change, CEO checks that it is reconciled in both; a report that contradicts its own document, or a peer's, is returned. Strike and correct in place — never delete a superseded line.
142
+ - Missions that build against an external spec carry `{mission}/spec-pins.json` (version + content hash). CEO requires CTO to pin before implementation and CQO to verify before PASS. Nothing is complete in the abstract.
143
+ - A registered convention/gotcha is linked from the index of **every role it names as an audience**. CEO treats an unreachable entry as an unregistered one.
144
+ - L role documents and all worker reports must carry `## Lessons Preflight` and a one-line `## Lessons Tally` (tally immediately before `## Implementation Notes`). `0 fired` is a valid tally; an omitted tally is not. CEO must reject reports that omit either section.
117
145
  - Every worker spawn declares its model explicitly — never the inherited CLI default. CEO requires the model in each CXX's Worker Task Briefs and Worker Evidence Manifest. **A silent or truncated worker is a rate limit until proven otherwise**: before treating a stalled round as a CXX failure, CEO asks for the usage-limit status and reset time, because a rate-limited worker looks exactly like a finished one from the outside.
118
146
  - Every worker report file is created **before the worker starts**, already carrying its required sections (seeded from `.harness/shared/templates/worker-report.md`). A worker killed mid-round must leave a valid partial report, never a stub. CEO treats a stub report as a seeding failure by the owning CXX, not as a worker failure.
119
147
  - Owner is the final acceptance reviewer, not a tester, QA substitute, debugger, or deployment verifier. CEO must not send "done, please check" reports while core functionality, regression, account setup, browser flows, logs, or runtime health remain unverified by workers.
@@ -121,7 +149,7 @@ The read is an **ordering constraint**, not a reading list. It happens before th
121
149
 
122
150
  ## Required Mission Note Format
123
151
 
124
- Every `ceo.md` and CXX document (`coo.md`, `cdo.md`, `cto.md`, `cqo.md`, `ops.md`) must end with this English section:
152
+ At L, every `ceo.md` and CXX document (`coo.md`, `cdo.md`, `cto.md`, `cqo.md`, `ops.md`) must end with this English section:
125
153
 
126
154
  ```
127
155
  ## Implementation Notes
@@ -146,10 +174,14 @@ Immediately above it, every document carries the one-line tally:
146
174
  - Fired: <items from Lessons Preflight that actually changed a decision, or `0 fired`>
147
175
  ```
148
176
 
149
- Use `None` when a subsection has no entries. These notes are mandatory even for small or emergency work. They must summarize how the role interpreted the Owner request, where the role intentionally diverged from the request, what alternatives were considered, and any true external-authority blocker. Do not list routine CEO-approved operations as "needs Owner confirmation."
177
+ Use `None` when a subsection has no entries. S/M role documents use the compact format above. They must summarize how the role interpreted the Owner request, where the role intentionally diverged from the request, what alternatives were considered, and any true external-authority blocker. Do not list routine CEO-approved operations as "needs Owner confirmation."
150
178
 
151
179
  When briefing a CXX, CEO must explicitly require the CXX to append this section to its own `{cxx}.md` and to require every worker it manages to append the same section to the bottom of that worker's report.
152
180
 
181
+ Prefer a genuinely separate CQO session (Claude Agent, Codex sub-agent, or `codex exec`) from implementation. Reading another role skill in the same session is only a role switch, not independent verification. If separate execution is unavailable, record `Verification Session: same-session` and disclose the limitation to Owner; the completion gate warns but permits it.
182
+
183
+ CEO discloses same-session verification and any termination without acceptance in the final Owner report.
184
+
153
185
  ## Routing Gate — CEO Must Never Bypass CXX
154
186
 
155
187
  CEO communicates **only** with CXX agents. CEO must **never**:
@@ -157,16 +189,18 @@ CEO communicates **only** with CXX agents. CEO must **never**:
157
189
  - Dispatch, hire, or brief specialist workers directly. Only CXX agents hire and manage workers.
158
190
  - Write documents on behalf of another CXX (i.e., author `cto.md`, `cqo.md`, `coo.md`, etc.). Each CXX owns its own document.
159
191
  - Mark a CXX step as complete without that CXX having run and produced its own document.
160
- - Skip a required CXX because the scope seems small. There is no scope exemption.
192
+ - Skip a required CXX because the scope seems small. Tier determines procedure depth; required implementation and verification roles remain separate.
161
193
 
162
194
  **Correct routing for every implementation mission:**
163
195
  ```
164
- Owner → CEO → CTO → [dev workers]
165
- └─── CQO → [evaluator/tester workers]
196
+ Owner → CEO → CTO → [dev workers] → CQO Handoff (tree frozen)
197
+ └→ CQO → [evaluator/tester workers] → verdict
166
198
  ```
167
199
 
168
- If CEO needs implementation done, CEO routes to CTO. CTO then hires dev workers.
169
- If CEO needs QA done, CEO routes to CQO. CQO then hires evaluator/tester workers.
200
+ CTO and CQO never run at the same time on the same tree. CEO routes to CQO only after CTO exits with a CQO Handoff; on FAIL, CEO routes back to CTO, and CQO resumes only after the next handoff.
201
+
202
+ If CEO needs implementation done, CEO routes to CTO. CTO implements directly at S, otherwise hires dev workers.
203
+ If CEO needs QA done, CEO routes to CQO. CQO directly runs tests at S, otherwise hires evaluator/tester workers.
170
204
  CEO does not contact workers. CXX contact workers.
171
205
 
172
206
  **When a CXX is unavailable or unresponsive:** retry with a fresh role-scoped context, route to another relevant CXX for recovery planning, or record an internal blocker with evidence. Escalate to the Owner only when the blocker requires external authority listed in the Autonomous Operating Charter.
@@ -94,6 +94,12 @@ Required output sections:
94
94
 
95
95
  Every COO worker brief must require the worker to append the same English `## Implementation Notes` block to the bottom of `.harness/documents/{mission_name}/coo/workers/{worker-name}.md`, covering risks, self-corrections, chosen direction, and unresolved questions. Use `None` for empty subsections.
96
96
 
97
+ ## The Document Is The Record
98
+
99
+ A conclusion you hold but have not written into `coo.md` **is not held by the company.** Before reporting any state change — to CEO, to a peer CXX, to the Owner — reconcile it in your own document *and* in `progress.json`. Strike and correct in place; never delete the superseded line, because a reader arriving later needs to see that it was superseded rather than never written.
100
+
101
+ Check your document against your peers' documents, not only against itself. The cheap version of this failure is a deliverable table that contradicts three messages you already sent. The expensive version was measured: a completed step reported and accepted, never written to the state file, and an orchestration loop that went on trying to spawn it **70 times**.
102
+
97
103
  ## Worker Spawn Contract
98
104
 
99
105
  Two things are decided **before** the round starts, not after a worker dies.
@@ -15,16 +15,35 @@ Before quality work, read `.harness/conventions/shared.md`, `.harness/convention
15
15
 
16
16
  ### Lessons Before Plan
17
17
 
18
+ At effective tier S/M, role documents may replace Lessons Preflight + Lessons Tally with `## Lessons` containing `Preflight: <applicable items and why>` (before work) and `Fired: <items or 0 fired>` (at completion). `## Implementation Notes` may contain concise bullets. At L, keep the full role format below. All worker reports retain the full seeded format at every tier.
19
+
18
20
  That read happens **before** the first source edit, the first measurement, and the first worker brief — not alongside them, and not after. The corpus is rarely the problem; the ordering is. Then, in `cqo.md`, write:
19
21
 
20
22
  - `## Lessons Preflight` — which convention/gotcha items apply to this mission and why, named by id or heading. Written before any worker is dispatched. If the corpus genuinely has nothing for this topic, say so explicitly.
21
23
  - `## Lessons Tally` — one line, written last, naming which of those items actually fired. **`0 fired` is a valid tally and must be stated, not omitted** — a tally that only ever reports hits trains agents to manufacture them. Place it immediately before `## Implementation Notes`.
22
24
 
23
- **Propagate verbatim.** Any requirement this skill places on CQO that its workers must also satisfy — the linked corpus items, the browser-automation clause, the seeded report skeleton, the `## Lessons Tally` line, the `## Implementation Notes` block — is copied **word for word** into every worker brief. *A rule stated one layer above the layer that executes it does not apply,* and a worker cannot infer a rule it was never given.
25
+ **Worker brief:** name the seeded report path and instruct the worker to fill its existing sections incrementally. Do not copy the report skeleton, Tally, or Notes block into the brief. Continue to pass relevant corpus links and copy behavioral requirements absent from the seed (including the browser-automation clause) verbatim.
24
26
 
25
27
  Do not distill the corpus into a private checklist file and read that instead. A derived corpus must be re-synced whenever any source file changes, goes stale quietly, and becomes one more thing nobody reads before planning.
26
28
 
27
- ## Workflow
29
+ ## Tier S: Direct Verification
30
+
31
+ At effective S, execute changed-scope tests and the available full suite directly (or cite a still-valid run under Evidence Reuse below), without hiring an evaluator. LLM-only inspection is not verification evidence. In cqo.md write `## Verification Commands` with each command, exit code, and output excerpt, and exactly one `Verification Session: separate|same-session` line (select the actual value). Also keep Lessons, Instrument Validity when relevant, OPS Watch Evidence, CQO Verdict, Recurrence Notes, and concise Implementation Notes.
32
+
33
+ Prefer a genuinely separate CQO session (Claude Agent, Codex sub-agent, or `codex exec`) from implementation. Reading another role skill in the same session is only a role switch, not independent verification. If separate execution is unavailable, record `Verification Session: same-session` and disclose the limitation to Owner; the completion gate warns but permits it.
34
+
35
+ OPS watch is required only when verification exercises a long-lived runtime (dev server, Docker, preview, cloud, device). A test runner or build that exits by itself is not one: write `OPS N/A: <reason>` in cqo.md and do not request OPS. At M, use one evaluator worker after CTO handoff; role Lessons/Notes may be compact. The worker workflow below applies to M/L.
36
+
37
+ ## Gate, Not Parallel
38
+
39
+ CQO is a pass gate, never a co-worker of CTO. A tree that is still being edited cannot be verified; every edit invalidates the fingerprint and restarts the suite.
40
+
41
+ 1. **Start condition.** CQO (and its evaluators) runs no test until `cto.md` has a `## CQO Handoff` (S: after `## Direct Work`) and CTO has exited with `agent_status="completed"`. Before that, CQO may only read context and write gate criteria in `cqo.md`.
42
+ 2. **Frozen tree.** If a run reports `tree changed during run` (exit 3) or the fingerprint moves while CQO is verifying, do not rerun. Record `Verdict: BLOCKED` with reason `tree-changed-during-verification` and route to CEO. CTO must finish and hand off again.
43
+ 3. **Fix loop.** After FAIL, CTO fixes and hands off again. Rerun only the failed tests plus the changed-scope tests for the new diff. The full suite runs once, on the final handed-off tree, immediately before PASS.
44
+ 4. **Parallel only in isolation.** If CQO must run concurrently for a real reason, it verifies a fixed worktree checkout of the handed-off commit, never the live tree.
45
+
46
+ ## Workflow (M/L)
28
47
 
29
48
  1. Read CEO and CTO mission context.
30
49
  2. Record decisions in `.harness/documents/{goal-or-child-mission}/cqo.md`.
@@ -33,7 +52,8 @@ Do not distill the corpus into a private checklist file and read that instead. A
33
52
  5. Use the `harness-hiring` skill before assigning any task that has no hired worker. Do not complete that task yourself.
34
53
  6. Define quality gates and delegate evidence collection to hired workers in fresh sessions.
35
54
  7. Monitor repeated issues and promote verified lessons to `.harness/conventions`, `.harness/gotchas`, `.harness/memories`, or `.harness/shared`.
36
- 8. Approve or reject archive based solely on worker-provided evidence and OPS runtime/watch evidence when the mission uses a runnable environment.
55
+ 8. Run `bash scripts/harness-spec-pin.sh . {goal-or-child-mission} verify` before any PASS and before archive. A category is complete **against a spec version**, never in the abstract; drift means the verified scope no longer matches what the work was built for, and the verdict is BLOCKED until CTO re-checks the affected work and re-pins.
56
+ 9. Approve or reject archive based solely on worker-provided evidence and OPS runtime/watch evidence when the mission uses a runnable environment.
37
57
 
38
58
  ## Worker Activity Telemetry
39
59
 
@@ -52,7 +72,7 @@ When the active goal is operating (perpetual, `mission-state.json` lifecycle `op
52
72
  CQO must verify the mission in two layers:
53
73
 
54
74
  1. Changed-scope verification: evaluator/tester workers inspect and test only the files, modules, APIs, flows, and adjacent dependencies identified in CTO's handoff.
55
- 2. Final full-suite gate: CQO runs the project's full test/coverage command once, near the end, through an evaluator/tester worker and normal project tooling.
75
+ 2. Final full-suite gate: CQO runs the project's full test/coverage command once, on the final handed-off tree (see Gate, Not Parallel); rerun only when Evidence Reuse conditions fail, through normal project tooling (directly at S, through an evaluator/tester worker at M/L).
56
76
 
57
77
  CQO must not ask evaluator workers to manually perform full-project test coverage analysis by LLM inspection. The full gate must use fast executable tooling such as `npm test`, `npm run test:coverage`, `pnpm test`, `pytest`, `go test ./...`, CI-equivalent scripts, or the repository's documented command. If no full-suite command exists, CQO records that as a verification gap instead of inventing a manual full-coverage review.
58
78
 
@@ -64,6 +84,38 @@ If changed-scope tests or changed-scope coverage fail, CQO returns FAIL or BLOCK
64
84
 
65
85
  CQO must report any out-of-scope full-suite failure or coverage deficit to CEO and CTO with command output, affected paths, and the classification above. CQO must not expand the mission into broad unrelated test-writing work unless CEO explicitly routes that as a new task.
66
86
 
87
+ ## Evidence Reuse
88
+
89
+ These rules apply at S/M/L. A source fingerprint is necessary, not sufficient, for reuse.
90
+
91
+ 1. Run full-suite and E2E commands through `bash scripts/harness-verify-fingerprint.sh run --log <ignored-log-path> [--root DIR] [--include FILE]... -- CMD...`. Copy its emitted `Baseline:` line verbatim into `## Verification Commands`, adding `runtime=` (e.g. `node -v`) and `criteria=` (acceptance-criteria reference). For E2E also record `target=` (host/device/account/build) and `preconditions=` (data/config/session). Commands must take credentials from environment variables, never literal arguments. Log only sanitized output. Relative log/include paths and CMD resolve from the repository root; keep logs gitignored and outside the include set.
92
+ 2. Reuse requires a retained log and script-produced `Baseline:` with `exit=0`; a fresh fingerprint command must exit 0 and match that baseline using the same `--include` inputs. `cmd` (including options), `runtime`, and `criteria` must also match. Record `Reused: <baseline document#item> fingerprint=<fp>` only after checking all conditions. A handwritten baseline or a run without a `Baseline:` line cannot be reused. Exit 3 alone is ambiguous: the command itself may return 3.
93
+ 3. Non-zero commands, including failures from pre-existing out-of-scope coverage thresholds, cannot supply reusable baselines. Use a command whose exit status matches the acceptance criteria (e.g. full tests separately from the existing changed-scope coverage gate); do not waive failed criteria.
94
+ 4. Include ignored environment input files and external E2E scripts with `--include`. Reinstalling dependencies or changing environment files invalidates the baseline. If unchanged dependencies/environment cannot be established, rerun the affected checks. E2E additionally requires identical verified target and preconditions; unknown live server data/session state requires rerunning it.
95
+ 5. A defect in verification logic (fail-open, unasserted controls, etc.) invalidates every piece of evidence relying on that logic.
96
+ 6. Wait until parallel product edits stop before recording a baseline, or verify in a fixed worktree checkout. Before/after hashes cannot detect a change that is reverted during execution (ABA). A changed tree emits no baseline and exits 3.
97
+ 7. At M/L, the evaluator worker decides reuse and CQO cites that report. Reuse has no fixed execution-count cap: valid evidence for the final tree is required, and unchanged conditions do not require another run.
98
+
99
+ ## Verification Artifact Hygiene
100
+
101
+ Pass this section and Evidence Reuse verbatim to workers writing or delegating verification scripts.
102
+
103
+ - **Existing tools first:** use the project's scenario runner, test runner, or Playwright configuration before writing a new script.
104
+ - **Secrets:** read credentials only from environment variables and fail non-zero before opening a browser if they are missing. Never put secret values in command arguments, logs, tool output, temporary files, or pattern files. Scan with `bash scripts/harness-secret-scan.sh <VAR_NAME>... -- <artifact-path>...`, never `grep "$VAR"`. Only exit 0 together with `leaks=0` means no leak; exit 2 is an incomplete scan, not a clean result. Do not enable shell tracing for secret handling.
105
+ - **Sanitization:** record URL paths without query strings. Do not record authentication headers, cookies, tickets, or raw request/response bodies; retain only safe key names, counts, and identifiers. Check forbidden patterns immediately before saving results and fail if found.
106
+ - **Captures:** capture only the element under verification or mask sensitive regions. Do not take full-screen captures containing personal information or faces.
107
+ - **Fail closed:** assert required steps instead of hiding them behind `if`; assert exact request counts rather than trusting `every()` on an empty array. Assert both positive and negative controls. Exceptions must exit non-zero; set `pass=true` only after every required check succeeds.
108
+ - **Prove failure paths:** record non-zero exits for missing credentials, missing target or zero requests, failed controls, and exceptions.
109
+ - **Retention:** keep artifacts in gitignored paths, preserve sanitized evidence referenced by a baseline, and delete unsanitized intermediate artifacts after judgment.
110
+
111
+ ## Reachability, Not Just Reading
112
+
113
+ Lazy loading is a **promise about reachability**. When you register a convention or gotcha, declare every role that should be able to find it — `<!-- roles: cto, cqo -->` at the top of a topic file, or `- **Roles**: cto, cqo` inside an index entry — and link it from **each** of those roles' index files, not only your own.
114
+
115
+ **The failure is filing under yourself.** Registration feels complete because the entry is indexed; it just is not where its declared readers are told to look. Measured on a live corpus: 69 items, **10 unreachable role-routings, 7 of them invisible to a role the entry itself named.** An agent that follows the reading rule exactly still never sees them — the rule stops narrowing the search and starts hiding the entry.
116
+
117
+ Verify with `bash scripts/harness-corpus-reachability.sh . text` (add `--fix` to link what is missing). This runs at the completion gate, so an unreachable corpus blocks the mission from closing.
118
+
67
119
  ## Instrument Validity
68
120
 
69
121
  Evidence about what did **not** happen is worth exactly as much as the instrument that looked for it.
@@ -72,21 +124,22 @@ Evidence about what did **not** happen is worth exactly as much as the instrumen
72
124
  - Report the control next to the result: what was injected, that it was observed, and the negative result from the same run. A verdict resting on unproven negative evidence is **BLOCKED**, not PASS.
73
125
  - **Where an instrument is supplied by a dependency rather than written in-repo, its filtering behaviour is read from source and quoted** — package, version, file, line range — not inferred from observed output. A filter that lives upstream is invisible to every in-repo search, so its absence from the project's own code is not evidence of its absence.
74
126
  - When an instrument turns out to have been structurally null, the claims it produced are identifiable **by their shape** — every claim of that form, not just the one that happened to be noticed. Re-open them as a class and say so in Recurrence Notes.
127
+ - A summary statistic is published **with its `n`**, and a spiky series is characterised by **percentiles, never min–max** — a range is the two least representative points in the set, and reads as a finding.
75
128
  - An audit question that offers alternatives asserts that the alternatives are exhaustive. "Is it A or B?" cannot return "neither, it is upstream". When an audit stalls, re-ask the question without the menu.
76
129
 
77
130
  ## Hard Rules
78
131
 
79
- CQO must not directly execute QA, visual review, security review, performance testing, or regression checks. CQO may only define gates, select evaluators, review evidence, decide archive eligibility, and document worker names and report paths.
132
+ At M/L, CQO must not directly execute QA, visual review, security review, performance testing, or regression checks. CQO may only define gates, select evaluators, review evidence, decide archive eligibility, and document worker names and report paths.
80
133
 
81
- **A verdict with no Worker Evidence Manifest is invalid.** CQO cannot issue ACCEPTED or REJECTED without at least one evaluator/tester worker record in `cqo.md`. Self-verification by CQO — where CQO writes a verdict based on its own inspection rather than worker-provided evidence — is a protocol violation. If no evaluator workers exist, use `harness-hiring` first.
134
+ **At M/L, a verdict with no Worker Evidence Manifest is invalid.** At M/L, CQO cannot issue ACCEPTED or REJECTED without at least one evaluator/tester worker record in `cqo.md`. LLM-only inspection is invalid at every tier; direct executed tests are permitted only at S. If no evaluator workers exist, use `harness-hiring` first.
82
135
 
83
136
  **CQO does not communicate with dev workers.** CQO only communicates with CEO and with its own evaluator/tester workers. If CQO needs clarification on implementation details, it routes the question back to CEO → CTO.
84
137
 
85
138
  Every evaluator/tester dispatched by CQO must write its report under `.harness/documents/{goal-or-child-mission}/cqo/workers/{worker-name}.md`.
86
139
 
87
- **Owner is not the QA tester.** CQO must not approve a handoff that asks the Owner to verify basic functionality, regression safety, browser behavior, account setup, logs, or runtime health. CQO must use evaluator/tester workers to collect the evidence, including E2E/Playwright/browser checks, regression commands, test-account or seeded-data validation, screenshots, logs, and risk notes when relevant. If evidence is missing, CQO verdict is BLOCKED or FAIL, not "ask Owner to check."
140
+ **Owner is not the QA tester.** CQO must not approve a handoff that asks the Owner to verify basic functionality, regression safety, browser behavior, account setup, logs, or runtime health. CQO must collect executed evidence directly at S or through evaluator/tester workers at M/L, including E2E/Playwright/browser checks, regression commands, test-account or seeded-data validation, screenshots, logs, and risk notes when relevant. If evidence is missing, CQO verdict is BLOCKED or FAIL, not "ask Owner to check."
88
141
 
89
- **OPS must watch runnable verification.** When CQO evaluator workers run Playwright, E2E, API, visual, accessibility, performance, or regression checks against a local/dev/preview/Docker/cloud runtime, CQO must request OPS monitoring before issuing PASS. CQO must include OPS evidence in `cqo.md` or mark the verdict BLOCKED. A CQO PASS is invalid if OPS reports an open INCIDENT, missing runtime mapping, required log missing, service down, health mismatch, or unmonitored runtime that is part of the tested scenario.
142
+ **OPS must watch long-lived runtime verification.** When CQO evaluator workers run Playwright, E2E, API, visual, accessibility, performance, or regression checks against a running local/dev/preview/Docker/cloud service (not a self-exiting test runner), CQO must request OPS monitoring before issuing PASS. CQO must include OPS evidence in `cqo.md` or mark the verdict BLOCKED. A CQO PASS is invalid if OPS reports an open INCIDENT, missing runtime mapping, required log missing, service down, health mismatch, or unmonitored runtime that is part of the tested scenario.
90
143
 
91
144
  If OPS reports an incident during verification:
92
145
 
@@ -101,33 +154,21 @@ Required output sections in `cqo.md`:
101
154
  2. Worker Task Briefs — gate, capability needed, selected evaluator or hiring request, declared model, acceptance criteria.
102
155
  3. Worker Evidence Manifest — worker name, declared model, report path, command or artifact evidence, status.
103
156
  4. Instrument Validity — for every negative claim: the instrument, its log level and filter (quoted from source when the instrument comes from a dependency), and the positive control that fired in the same run. Negative evidence with no control is BLOCKED, not PASS.
104
- 5. OPS Watch Evidence — ops report path, monitored runtime mapping, incidents/warnings, and whether runtime evidence permits PASS.
105
- 6. CQO Verdict — PASS, FAIL, or BLOCKED based only on worker evidence plus required OPS watch evidence. Must reference Worker Evidence Manifest entries.
106
- 7. Recurrence Notes — accepted gotchas, conventions, memories, or none.
157
+ 5. OPS Watch Evidence — ops report path, monitored runtime mapping, incidents/warnings, and whether runtime evidence permits PASS; or one line `OPS N/A: <reason>` when no long-lived runtime was tested.
158
+ 6. CQO Verdict — inside `## CQO Verdict`, write a dedicated `Verdict: PASS|ACCEPTED|FAIL|REJECTED|BLOCKED` line with exactly one uppercase value. The last verdict candidate across these sections decides: trailing commentary, empty value, bold decoration, or lowercase makes it invalid; no fallback to an older PASS. Use the canonical `## CQO Verdict` heading for re-evaluation too. The completion reader also includes nested headings and parenthesized headings such as `## CQO Verdict (Re-test)` until the next unrelated level-1/2 heading. Label decoration (`**Verdict**: FAIL`) or spacing before the colon (`Verdict : FAIL`) still counts as an attempted verdict but is rejected as invalid. Put reasons on the next line. Correct old verdicts using `~~Verdict: FAIL~~` and a new line. Cite executed evidence (worker manifest at M/L) plus required OPS evidence.
159
+ 7. Recurrence Notes — accepted gotchas, conventions, memories, or `none — <reason>` at S/M when recurrence is unlikely and the cause is obvious (L hot-fixes still register a lesson). Every entry registered here names **every role that should be able to find it** and is linked from each of those roles' indexes; `scripts/harness-corpus-reachability.sh` must pass.
107
160
  8. Lessons Tally — one line naming which preflight items actually fired. `0 fired` is valid and must be stated.
108
161
  9. Implementation Notes — in English, with `Design Decisions`, `Deviations`, `Tradeoffs`, and `Open Questions`.
109
162
 
110
163
  ## Worker Report Note Requirement
111
164
 
112
- Every CQO evaluator/tester brief must require the worker to append this English block to the bottom of `.harness/documents/{goal-or-child-mission}/cqo/workers/{worker-name}.md`:
113
-
114
- ```
115
- ## Implementation Notes
116
-
117
- ### Design Decisions
118
- - ...
119
-
120
- ### Deviations
121
- - ...
165
+ Point the worker to its seeded report path. Require it to fill the existing Implementation Notes (all four subsections, `None` when empty); do not duplicate the template in the brief.
122
166
 
123
- ### Tradeoffs
124
- - ...
167
+ ## The Document Is The Record
125
168
 
126
- ### Open Questions
127
- - ...
128
- ```
169
+ A conclusion you hold but have not written into `cqo.md` **is not held by the company.** Before reporting any state change — to CEO, to a peer CXX, to the Owner — reconcile it in your own document *and* in `progress.json`. Strike and correct in place; never delete the superseded line, because a reader arriving later needs to see that it was superseded rather than never written.
129
170
 
130
- The worker notes must cover risks, self-corrections, and chosen direction. Use `None` when a subsection has no entries. CQO must not accept evaluator output that omits this block.
171
+ Check the role document against peer documents and runtime state before reporting completion.
131
172
 
132
173
  ## Worker Spawn Contract
133
174
 
@@ -137,4 +178,4 @@ Two things are decided **before** the round starts, not after a worker dies.
137
178
 
138
179
  **2. Seed the report.** Create `.harness/documents/{goal-or-child-mission}/cqo/workers/{worker-name}.md` **before the worker starts**, already carrying every required section — `## Status` (`IN_PROGRESS`), `## Task`, `## Evidence`, `## Result`, `## Lessons Tally`, and the terminal `## Implementation Notes` block with all four subsections stubbed. Copy `.harness/shared/templates/worker-report.md` when it is installed; otherwise write the skeleton by hand. Brief the worker to fill it in **incrementally as the work happens**, never to assemble the report at the end.
139
180
 
140
- A worker that dies mid-round — rate limit, crash, cancelled session — must leave a **valid partial report, never a stub**. Same failure, opposite outcome, one variable: unseeded workers killed mid-round left stubs and halted the company; a seeded worker killed by the same limit left its report intact and cost nothing. The variable was a decision taken before the round.
181
+ A worker interrupted mid-round must leave a valid partial report, never a stub.
@@ -15,16 +15,24 @@ Before engineering work, read `.harness/conventions/shared.md`, `.harness/conven
15
15
 
16
16
  ### Lessons Before Plan
17
17
 
18
+ At effective tier S/M, role documents may replace Lessons Preflight + Lessons Tally with `## Lessons` containing `Preflight: <applicable items and why>` (before work) and `Fired: <items or 0 fired>` (at completion). `## Implementation Notes` may contain concise bullets. At L, keep the full role format below. All worker reports retain the full seeded format at every tier.
19
+
18
20
  That read happens **before** the first source edit, the first measurement, and the first worker brief — not alongside them, and not after. The corpus is rarely the problem; the ordering is. Then, in `cto.md`, write:
19
21
 
20
22
  - `## Lessons Preflight` — which convention/gotcha items apply to this mission and why, named by id or heading. Written before any worker is dispatched. If the corpus genuinely has nothing for this topic, say so explicitly.
21
23
  - `## Lessons Tally` — one line, written last, naming which of those items actually fired. **`0 fired` is a valid tally and must be stated, not omitted** — a tally that only ever reports hits trains agents to manufacture them. Place it immediately before `## Implementation Notes`.
22
24
 
23
- **Propagate verbatim.** Any requirement this skill places on CTO that its workers must also satisfy — the linked corpus items, the browser-automation clause, the seeded report skeleton, the `## Lessons Tally` line, the `## Implementation Notes` block — is copied **word for word** into every worker brief. *A rule stated one layer above the layer that executes it does not apply,* and a worker cannot infer a rule it was never given.
25
+ **Worker brief:** name the seeded report path and instruct the worker to fill its existing sections incrementally. Do not copy the report skeleton, Tally, or Notes block into the brief. Continue to pass relevant corpus links and copy behavioral requirements absent from the seed (including the browser-automation clause) verbatim.
24
26
 
25
27
  Do not distill the corpus into a private checklist file and read that instead. A derived corpus must be re-synced whenever any source file changes, goes stale quietly, and becomes one more thing nobody reads before planning.
26
28
 
27
- ## Workflow
29
+ ## Tier S: Direct Work
30
+
31
+ At effective S, create cto.md before edits, implement directly without hiring/worker telemetry, and run changed-scope tests. Record `## Direct Work` (changes and executed evidence), `## CQO Handoff` (scope, commands, risks, available full-suite command), `## Lessons`, and concise `## Implementation Notes`. Hand off to a separate CQO session; do not issue your own quality verdict. The worker workflow below applies to M/L.
32
+
33
+ On S→M/L, freeze Direct Work using the CEO-recorded SHA. Add `## Post-Upgrade Work`; write `none — <reason>` first if no added implementation, otherwise delegate it to CTO workers. Final CQO verification follows the higher tier. M keeps Worker Task Briefs/Manifest and CQO Handoff but uses compact Lessons/Notes; L uses the full output below.
34
+
35
+ ## Workflow (M/L)
28
36
 
29
37
  1. Read CEO, COO, and CDO mission documents.
30
38
  2. Record decisions in `.harness/documents/{goal-or-child-mission}/cto.md`. **This file must be created before any worker is dispatched.**
@@ -57,8 +65,32 @@ Every CTO worker brief that may use Playwright, browser automation, browser-base
57
65
 
58
66
  CTO must not accept worker plans or reports that omit this requirement when browser automation is in scope.
59
67
 
68
+ ## Reachability, Not Just Reading
69
+
70
+ Lazy loading is a **promise about reachability**. When you register a convention or gotcha, declare every role that should be able to find it — `<!-- roles: cto, cqo -->` at the top of a topic file, or `- **Roles**: cto, cqo` inside an index entry — and link it from **each** of those roles' index files, not only your own.
71
+
72
+ **The failure is filing under yourself.** Registration feels complete because the entry is indexed; it just is not where its declared readers are told to look. Measured on a live corpus: 69 items, **10 unreachable role-routings, 7 of them invisible to a role the entry itself named.** An agent that follows the reading rule exactly still never sees them — the rule stops narrowing the search and starts hiding the entry.
73
+
74
+ Verify with `bash scripts/harness-corpus-reachability.sh . text` (add `--fix` to link what is missing). This runs at the completion gate, so an unreachable corpus blocks the mission from closing.
75
+
76
+ ## Spec Version Pins
77
+
78
+ Nothing is complete in the abstract. **A category is complete against a spec version.**
79
+
80
+ When the mission builds against any external spec — an API contract, a schema, a partner document, a standard — record it before implementation starts:
81
+
82
+ ```
83
+ bash scripts/harness-spec-pin.sh . {goal-or-child-mission} add <name> <path> <version>
84
+ ```
85
+
86
+ This stores the version **and a content hash** in `{mission}/spec-pins.json`. Re-run `... verify` before handing off to CQO and before declaring any category done. The completion gate re-checks the pins, so drift blocks the mission from closing.
87
+
88
+ Measured: a spec moved `v0.7 → v0.9`, changing a response contract, while the category built against `v0.7` sat marked **complete**. Two later revisions had landed silently. The symptom was not an error — a lookup key stopped matching, and three overlays were dropped as `null`. **No error, no log, just an absence.** Nothing in the harness recorded which version the work had been for.
89
+
60
90
  ## Test Coverage Scope
61
91
 
92
+ When writing or delegating verification scripts, follow CQO’s Verification Artifact Hygiene and Evidence Reuse sections and copy both sections verbatim into the worker brief.
93
+
62
94
  CTO must optimize engineering verification around the work actually changed in the mission. CTO worker briefs must require targeted tests, coverage checks, and regression commands for the changed files, modules, APIs, flows, and directly affected dependencies only.
63
95
 
64
96
  CTO must not require workers to manually reason through full-project test coverage, inspect unrelated coverage gaps, or chase 100% coverage outside the modified scope. Full-project test execution belongs to CQO's final gate and must be run by project test tooling, not by LLM inspection.
@@ -71,11 +103,13 @@ When handing off to CQO, CTO must include:
71
103
  - Known risk areas and directly adjacent dependencies.
72
104
  - Suggested full-suite command if the project exposes one, such as `npm test`, `npm run test:coverage`, `pnpm test`, `pytest`, `go test ./...`, or the repo's equivalent.
73
105
 
106
+ **Handoff freezes the tree.** Write `## CQO Handoff` only when implementation is finished, then exit with `agent_status="completed"`. From that point CTO and its workers make no edits until CQO returns a verdict; a needed change goes through CEO as a new fix iteration after the verdict. CQO does not start before this handoff.
107
+
74
108
  If targeted verification fails inside the changed scope, CTO blocks the handoff until workers fix or explicitly document the blocker. If unrelated tests or coverage gaps are noticed outside the changed scope, CTO records them as possible pre-existing risk or side-effect signal and routes them through CEO/CQO instead of expanding the implementation mission by default.
75
109
 
76
110
  ## Hard Rules
77
111
 
78
- CTO must not directly write code, create build scripts, choose detailed implementation content, run technical QA as the evaluator, or produce final implementation artifacts. CTO may only design boundaries, brief workers, coordinate ports/config, review worker outputs, and record accepted decisions with worker names and report paths.
112
+ At M/L, CTO must not directly write code, create build scripts, choose detailed implementation content, run technical QA as the evaluator, or produce final implementation artifacts. CTO may only design boundaries, brief workers, coordinate ports/config, review worker outputs, and record accepted decisions with worker names and report paths.
79
113
 
80
114
  **cto.md is a prerequisite gate.** No worker may be dispatched before `cto.md` exists. A mission where workers appear in `.harness/documents/{goal-or-child-mission}/cto/workers/` but no `cto.md` exists is a protocol violation — CEO bypassed CTO.
81
115
 
@@ -83,7 +117,7 @@ Every worker dispatched by CTO must be listed in the Worker Evidence Manifest se
83
117
 
84
118
  **Owner is not the technical tester.** CTO must not hand unfinished software to CEO/Owner with "please check" as the validation plan. CTO must require workers to prove implementation readiness with appropriate unit tests, integration checks, build/run commands, seeded data or test account setup, browser/E2E checks when applicable, and changed-file evidence. If verification cannot be completed, CTO reports BLOCKED with the missing evidence instead of asking the Owner to test it.
85
119
 
86
- Required output sections in `cto.md`:
120
+ Required output sections in `cto.md` at L (M uses compact Lessons/Notes; S uses Direct Work above):
87
121
 
88
122
  1. Lessons Preflight — convention/gotcha items that apply to this mission, why each applies, and the topic links passed into worker briefs. Written before the first worker is dispatched.
89
123
  2. Worker Task Briefs — task, capability needed, selected worker or hiring request, declared model, acceptance criteria.
@@ -96,25 +130,13 @@ Required output sections in `cto.md`:
96
130
 
97
131
  ## Worker Report Note Requirement
98
132
 
99
- Every CTO worker brief must require the worker to append this English block to the bottom of `.harness/documents/{goal-or-child-mission}/cto/workers/{worker-name}.md`:
133
+ Point the worker to its seeded report path. Require it to fill the existing Implementation Notes (all four subsections, `None` when empty); do not duplicate the template in the brief.
100
134
 
101
- ```
102
- ## Implementation Notes
103
-
104
- ### Design Decisions
105
- - ...
106
-
107
- ### Deviations
108
- - ...
135
+ ## The Document Is The Record
109
136
 
110
- ### Tradeoffs
111
- - ...
112
-
113
- ### Open Questions
114
- - ...
115
- ```
137
+ A conclusion you hold but have not written into `cto.md` **is not held by the company.** Before reporting any state change — to CEO, to a peer CXX, to the Owner — reconcile it in your own document *and* in `progress.json`. Strike and correct in place; never delete the superseded line, because a reader arriving later needs to see that it was superseded rather than never written.
116
138
 
117
- The worker notes must cover risks, self-corrections, and chosen direction. Use `None` when a subsection has no entries. CTO must not accept worker output that omits this block.
139
+ Check the role document against peer documents and runtime state before reporting completion.
118
140
 
119
141
  ## Worker Spawn Contract
120
142
 
@@ -124,4 +146,4 @@ Two things are decided **before** the round starts, not after a worker dies.
124
146
 
125
147
  **2. Seed the report.** Create `.harness/documents/{goal-or-child-mission}/cto/workers/{worker-name}.md` **before the worker starts**, already carrying every required section — `## Status` (`IN_PROGRESS`), `## Task`, `## Evidence`, `## Result`, `## Lessons Tally`, and the terminal `## Implementation Notes` block with all four subsections stubbed. Copy `.harness/shared/templates/worker-report.md` when it is installed; otherwise write the skeleton by hand. Brief the worker to fill it in **incrementally as the work happens**, never to assemble the report at the end.
126
148
 
127
- A worker that dies mid-round — rate limit, crash, cancelled session — must leave a **valid partial report, never a stub**. Same failure, opposite outcome, one variable: unseeded workers killed mid-round left stubs and halted the company; a seeded worker killed by the same limit left its report intact and cost nothing. The variable was a decision taken before the round.
149
+ A worker interrupted mid-round must leave a valid partial report, never a stub.