token-harness 0.1.9 → 0.1.11

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (4) hide show
  1. package/README.md +385 -266
  2. package/package.json +1 -1
  3. package/sbom.json +3 -3
  4. package/token-harness.mjs +30387 -20033
package/README.md CHANGED
@@ -1,405 +1,524 @@
1
1
  # Token Harness
2
2
 
3
- **Make Claude Code and Codex easier to understand and use efficiently.**
3
+ **Build, verify and measure an optimization stack for Claude Code and Codex.**
4
4
 
5
- Token Harness checks your coding agents, shows subscription allowance when it can be
6
- observed reliably, reduces avoidable context overhead, and recommends useful actions.
7
- It runs locally and never presents local token estimates as subscription quota.
5
+ Token Harness is a local **optimization stack manager**. It checks your coding agents, manages the
6
+ optimization components it can safely own, keeps experimental candidates separate, verifies the
7
+ result, and reports savings only when it has evidence to support them.
8
8
 
9
- ## Open it. Approve setup. Keep coding.
9
+ It is not another coding agent and it does not replace specialized projects such as RTK or
10
+ HarnessTrim.
10
11
 
11
- You need [Node.js 22.13 or newer](https://nodejs.org/) and an installed, signed-in
12
- Claude Code or Codex. Install Token Harness, then open it:
12
+ ## Start here
13
+
14
+ Requirements:
15
+
16
+ - Node.js 22.13 or newer;
17
+ - Claude Code or Codex installed;
18
+ - the coding agent you want to use already signed in.
19
+
20
+ Install and open Token Harness:
13
21
 
14
22
  ```sh
15
23
  npm install --global token-harness@latest
16
24
  token-harness
17
25
  ```
18
26
 
19
- The browser is now the primary interface. **Set up automatically** checks both agents and
20
- prepares the supported integration changes. It describes each change in plain language.
21
- Choose **Approve and apply** to apply the reviewed configuration with backups and verification.
22
- There are no plan IDs to copy and no daily command sequence to remember.
27
+ That is the normal human workflow. The browser app is the primary interface; there is no daily list
28
+ of CLI commands to memorize.
29
+
30
+ ### First run
31
+
32
+ 1. Open **Dashboard** and let Token Harness inspect the current setup.
33
+ 2. If setup is incomplete, choose **Open setup**.
34
+ 3. In **Setup**, work from top to bottom:
35
+ - Coding agents
36
+ - Managed optimizers
37
+ - Experimental tools
38
+ - Optional agent tuning
39
+ - Checks and maintenance
40
+ 4. For Claude Code or Codex, choose **Review setup** when a managed setup is available.
41
+ 5. Read the exact proposed changes. **Apply reviewed setup** appears only when there is a concrete
42
+ safe plan to apply.
43
+ 6. Keep using Claude Code or Codex normally.
44
+ 7. Open **Results** when you want to see what Token Harness can actually prove.
45
+
46
+ Opening the app does not change your configuration. Read-only checks stay read-only, and a managed
47
+ write requires an explicit review and approval.
23
48
 
24
- Already configured? The app shows your existing integrations without replacing them.
25
- An absent provider or an unreviewed version combination is explained rather than installed
26
- or forced silently. Automatic setup covers the reviewed integration paths, not every possible
27
- provider/version. Token Harness does not install Claude Code or Codex or log you in.
49
+ ## The three views
50
+
51
+ ### Dashboard
52
+
53
+ Dashboard answers the questions that matter first:
54
+
55
+ - is my setup ready;
56
+ - which coding agents and managed optimizers are active;
57
+ - what should I do next;
58
+ - what value has actually been measured;
59
+ - whether quality or an integration needs attention.
28
60
 
29
- ### Daily use
61
+ The headline cards deliberately distinguish measured evidence from unknown values. Missing evidence
62
+ is never shown as zero savings.
30
63
 
31
- Continue launching `claude` or `codex` as usual. Supported output integrations operate in the
32
- agent, not in the dashboard. You can close the page and its terminal without disabling those
33
- integrations. Open `token-harness` whenever you want to see results; the visible dashboard
34
- imports available provider records and refreshes its readings automatically.
64
+ ### Setup
65
+
66
+ Setup is one ordered workflow instead of a collection of unrelated actions.
35
67
 
36
- **Recorded savings** shows retained history across locally recorded projects, with date bounds,
37
- provider, measurement class, units, changed-output counts, and before/after values. It does not
38
- add incompatible provider figures together. Negative results remain visible. No telemetry is
39
- shown as **not measured**, never a reassuring zero or an invented subscription saving.
40
- Some provider records may predate Token Harness; locally stored records are not guaranteed
41
- complete lifetime history. RTK history is imported directly. HarnessTrim project-local records
42
- must have been imported from their project, or exposed through a configured known metrics path;
43
- the app does not crawl your disk looking for private projects.
68
+ **1. Coding agents**
44
69
 
45
- For a terminal-only summary, the one command is:
70
+ Token Harness currently supports guided setup for Claude Code and Codex. Agent details also show
71
+ useful read-only allowance and connected-tool observations when available.
46
72
 
47
- ```sh
48
- token-harness savings
49
- ```
73
+ **2. Managed optimizers**
50
74
 
51
- Optional windows are `--since 7d` and `--since 30d`. The advanced `metrics` command remains
52
- project-scoped; opening the app from its installation folder does not change the savings scope.
75
+ The managed stack currently consists of **RTK + HarnessTrim** on individually reviewed
76
+ combinations. Token Harness tracks the exact combined provider set separately: if no combined-stack
77
+ review is recorded, Setup says so and keeps the stack incomplete rather than inferring compatibility
78
+ from healthy individual checks. Token Harness can prepare their integration transactionally, show the
79
+ exact plan, apply it only after approval, verify it, and remove only configuration it owns.
53
80
 
54
- ### Find what you need
81
+ For maintainers validating the combined stack, `token-harness stack-review --json` captures the exact
82
+ configured provider versions and managed harness sets and reuses the existing passive `verify`
83
+ evidence for exact provider/harness pairs. Runtime evidence is credited only when it can be
84
+ attributed to that harness: HarnessTrim uses its native event harness field, while provider-wide
85
+ telemetry can be attributed by exclusion only when one harness is wired. With multiple harnesses, an
86
+ unattributable receipt stays **Unavailable** instead of being copied across rows. A `not-exercised`
87
+ result does not promise that ordinary agent use will create a receipt: the provider must record a
88
+ qualifying operation, and for reducers that means a real reduction. `stack-review` does not run an
89
+ active canary or spend a model call, and it never makes the compatibility decision itself. The
90
+ shipped combined-review registry stays empty until the captured configuration has enough real
91
+ verification evidence, has been benchmarked together, and has been deliberately reviewed; see
92
+ `docs/combined-stack-reviews.md`.
55
93
 
56
- The app has three tabs: **Overview** for agents and recorded savings, **Rules & settings**
57
- for one agent's rules at a time, and **Activity** for checks and guarded undo. Theme follows
58
- your system; the header also offers light and dark modes.
94
+ **3. Experimental tools**
59
95
 
60
- Each rule shows what was actually **observed**, separately from how it works. Actions sit
61
- next to the relevant state: **Adjust reasoning** opens the supported review/apply flow;
62
- **Change in Claude/Codex** explains native steps when an automatic write is not available.
63
- Missing allowance data and measurement records have their own setup/help actions.
96
+ Headroom, mcptoon and GitNexus are visible as candidates, not silently promoted dependencies. Their
97
+ cards distinguish CLI installation, candidate-side activation/evaluation, Token Harness benchmark
98
+ evidence and promotion-readiness gates. External install or activation commands are shown for
99
+ review; Token Harness does not silently execute package managers, activate wrappers, index
100
+ repositories or register MCP servers.
64
101
 
65
- A saved Claude effort can be displayed even on an unreviewed CLI version. That does not
66
- admit automatic writes on that version. **No saved preference** means the user field is
67
- absent, not that reasoning is disabled. A failed read is a separate state with a cause.
68
- The displayed preference is not a live reading of an already-running session.
102
+ Choose **Run standard evaluation** to start or resume a paired candidate campaign. Campaign state is
103
+ scoped to both the candidate and the selected harness, so Claude Code and Codex evidence cannot be
104
+ mixed accidentally. The browser reads the campaign directly and shows **Progress**, the current
105
+ selection signal, whether the evidence is **Decision ready**, the number of evidence-bearing pairs
106
+ and the exact **Next** step.
69
107
 
70
- ### The rules are visible
108
+ For normal use, the browser can now start and finish the local benchmark capture itself. You still
109
+ run the actual task in Claude Code or Codex. When the task finishes, record the quality result,
110
+ attempt count and failed-attempt count you actually observed. Before an optimized capture, Token
111
+ Harness requires you to acknowledge that you enabled the candidate through its own documented
112
+ workflow. That acknowledgement is **not activation verification** and is never treated as promotion
113
+ evidence. Every browser capture action is matched against the campaign engine's current step before
114
+ it can write local benchmark state, so a stale tab cannot advance a different step.
71
115
 
72
- **Rules & settings** explains each configured rule: what it does, why it is used,
73
- its mode, and the evidence available. Automatic integrations, persistent preferences,
74
- observations and features that are not enabled are explicitly distinguished.
116
+ `benchmark-matrix`, `benchmark-start` and `benchmark-finish` remain available as advanced terminal
117
+ fallbacks for debugging or automation. The browser does not run the coding task, install or activate
118
+ a candidate, or claim that candidate attribution proves activation.
75
119
 
76
- RTK's supported command integration can reduce output automatically. HarnessTrim can use
77
- adapters or skills/instructions, depending on the installation; skills-only is not a transparent
78
- hook, and the agent must actually use the reducer. Configured never means every command was
79
- intercepted. Provider telemetry and its exact/estimated classification remain separate evidence.
120
+ A campaign selection assessment can report `insufficient-evidence`, `promising`, `mixed` or
121
+ `negative`, plus whether the evidence is decision-ready. **Decision-ready is not promotion-ready.**
122
+ The campaign shows the same promotion-review gate count and next gate as the candidate card, using
123
+ the current local observation rather than starting another environment scan. Activation
124
+ verification, managed lifecycle, compatibility/reversibility, project maturity and combined-stack
125
+ validation remain separate gates.
80
126
 
81
- **Optional: match reasoning to your work** lets you choose the agent and the type of work
82
- without learning CLI flags. It previews the actual supported effort/verbosity change, then
83
- applies only after approval. **This is a persistent preference for future sessions, not an
84
- automatic per-task switch.** It does not switch models, billing, login, or hook trust. The
85
- baseline automatic setup never guesses a task or quietly lowers reasoning.
127
+ Choose **Compare evaluation evidence** when you want one read-only view of campaigns you already
128
+ started. The comparison is loaded only on request, never creates a campaign, keeps candidate and
129
+ harness ordering fixed, and reports progress, selection signal, decision readiness, runtime
130
+ activation evidence and promotion gates without producing a composite score or automatic winner.
131
+ Activation is shown separately as **Verified**, **Blocked** or **Unreviewed**, with
132
+ verified/blocked/unknown pair counts when campaign evidence exists.
86
133
 
87
- **Check integrations** performs the existing integration checks from the UI.
88
- **Undo last change**, available after an application in that dashboard session, previews a
89
- whole-file backup restoration. It refuses to undo a newer unrelated transaction. It restores
90
- only the last successful agent transaction; manual edits to those same files after that
91
- transaction would also be restored, as the confirmation explains.
134
+ **4. Optional agent tuning**
92
135
 
93
- ### Run the current source
136
+ Reasoning preferences are separate from optimizer installation. They are persistent agent settings,
137
+ not hidden per-task switches, and are changed only through the normal preview/apply flow.
94
138
 
95
- From an existing clone, after installing its dependencies, one command builds and opens the app:
139
+ **5. Checks and maintenance**
96
140
 
97
- ```sh
98
- npm start
99
- ```
141
+ Read-only integration checks, update checks and safe removal/undo controls live here. An update
142
+ outside reviewed compatibility is not forced.
100
143
 
101
- For a fresh clone:
144
+ ### Results
102
145
 
103
- ```sh
104
- git clone https://github.com/giuliastro/token-harness.git
105
- cd token-harness
106
- npx --yes pnpm@10.33.4 install --frozen-lockfile
107
- npm start
108
- ```
146
+ Results keeps evidence separate from estimates. Depending on what is actually observable, it can
147
+ show:
109
148
 
110
- This uses the clone, not an older global installation. An unmerged branch or unpublished main
111
- change is not automatically available through `token-harness@latest`.
149
+ - recorded optimizer output reduction;
150
+ - authoritative paired 5-hour / 7-day allowance evidence;
151
+ - API cost only when billed-token evidence and a verified price basis exist;
152
+ - paired quality evidence;
153
+ - experimental candidate evidence and campaign selection assessment when available;
154
+ - recent checks and changes from the current local app session.
112
155
 
113
- ### Advanced and AI-assisted use
156
+ **Not measured** means exactly that. Token Harness does not turn local token estimates into fake
157
+ subscription minutes, money or quota savings.
114
158
 
115
- An AI may use the existing JSON CLI to inspect, plan and apply an explicitly approved change.
116
- That is optional: another AI subscription is not required to operate the app. No persistent
117
- agent, background model calls or task classifier runs behind your back.
159
+ ## Daily use
118
160
 
119
- The older automation contracts remain available: `setup`, `optimize`, `plan`, `apply`,
120
- `verify`, `metrics`, `rollback`, and their JSON reports. `ui --json` preserves its existing
121
- schema-1 report; `ui --read-only` opens the legacy read-only dashboard. `ui --no-open` starts
122
- the guided app without launching a browser. Stop either local server with Ctrl+C.
161
+ Keep launching `claude` or `codex` as usual. Deterministic optimizers that are installed, verified
162
+ and still beneficial are intended to remain enabled.
123
163
 
124
- ## What normal output looks like
164
+ Token Harness does **not** need to stay open, does not need a permanent background daemon, and does
165
+ not need to decide before every command whether an optimizer should run.
125
166
 
126
- A healthy final check is intentionally short:
167
+ Open `token-harness` when you want to inspect health/results, review a setup change, check an update,
168
+ verify integrations or re-evaluate the stack after a meaningful version/configuration change.
127
169
 
128
- ```text
129
- TOKEN HARNESS - READY
170
+ The app does not periodically reload the whole setup. A full read happens on initial open or when you
171
+ choose **Refresh**. Existing readings remain visible while a refresh runs. Applying a reviewed change
172
+ marks the displayed data as previous state instead of immediately launching another expensive full
173
+ read.
130
174
 
131
- WHAT WORKS
132
- Codex: configured (0.146.0)
133
- HarnessTrim: active on Codex
175
+ ## What counts as savings
134
176
 
135
- CHANGES
136
- Nothing changed.
177
+ Token Harness keeps different evidence classes separate.
137
178
 
138
- NEXT STEP
139
- Use your coding agent normally; configured optimizers run automatically.
140
- ```
179
+ **Recorded output savings** are attributable reducer measurements. Providers, units and measurement
180
+ classes are not silently added together. Negative results and errors remain visible.
141
181
 
142
- A newer-than-tested combination is not presented as if the whole setup were broken:
182
+ **5h / 7d allowance savings** require authoritative paired before/after allowance evidence. A
183
+ five-hour percentage may also be expressed as the equivalent share of that 300-minute allowance
184
+ window. Weekly quota is not converted into seven days of wall-clock compute.
143
185
 
144
- ```text
145
- TOKEN HARNESS - READY WITH LIMITATIONS
186
+ **API cost** stays **Not measured yet** until attributable billed input/output tokens and a verified
187
+ model-price basis are available.
146
188
 
147
- WHAT WORKS
148
- Claude Code: configured
149
- RTK: active on Claude Code
189
+ **Quality** is measured independently. A measured regression blocks a positive allowance-saving
190
+ claim rather than letting a smaller token number win by itself.
150
191
 
151
- NEXT STEP
152
- token-harness verify
153
- You can keep working; verify the active integrations when convenient.
154
- ```
192
+ Upstream benchmark numbers are useful for deciding what to test; they are never copied directly into
193
+ your savings total.
155
194
 
156
- Need the evidence behind a summary? Add `--verbose`:
195
+ For a terminal-only savings summary:
157
196
 
158
197
  ```sh
159
- token-harness doctor --verbose
198
+ token-harness savings
160
199
  ```
161
200
 
162
- Need stable machine-readable output for automation? Add `--json`:
201
+ Optional windows are `--since 7d` and `--since 30d`.
163
202
 
164
- ```sh
165
- token-harness doctor --json
166
- token-harness ui --json
167
- ```
203
+ ## Current optimization stack
168
204
 
169
- `--json` keeps the complete schema-1 result and diagnostics; it is not shortened.
205
+ Token Harness prefers thin integrations around strong specialized projects instead of copying their
206
+ algorithms into this repository.
170
207
 
171
- ## Two entry points to remember
208
+ | Component | Role | Management |
209
+ | --- | --- | --- |
210
+ | [RTK](https://github.com/rtk-ai/rtk) | Shell/tool output reduction | Managed on reviewed combinations |
211
+ | [HarnessTrim](https://github.com/giuliastro/HarnessTrim) | Deterministic output/context reduction | Managed first-party integration |
212
+ | [cclimits](https://github.com/cruzanstx/cclimits) | Optional Claude allowance evidence | Read-only evidence; not an optimizer |
213
+ | [ccusage](https://github.com/ccusage/ccusage) | Local usage history | Read-only evidence; never subscription quota |
214
+
215
+ Provider compatibility is deliberately **not pinned forever to the first fixture version**. The
216
+ current compatibility policy includes RTK **0.49.0** (source-contract reviewed; the latest live
217
+ Windows harness-mutation fixture is 0.48.0) and HarnessTrim **0.3.0**. Newer HarnessTrim builds can
218
+ be accepted without another hard-coded version bump when their executable version matches their
219
+ machine-readable `capabilities` version and the semantic surface/write-set comparison reports no
220
+ drift.
221
+
222
+ Provider **package updates are separate from harness configuration writes**. `token-harness update`
223
+ can replace a reviewed provider target without requiring an exact historical Claude/Codex fixture
224
+ for that package version; exact compatibility rows still gate any later managed agent-config
225
+ mutation. HarnessTrim updates use its reviewed npm channel and capture the previous global version
226
+ for rollback. On native Windows RTK still prefers WinGet, but when that catalog is behind the
227
+ reviewed 0.49.0 target Token Harness can fall back to the exact official GitHub Windows x64 release:
228
+ it verifies GitHub's published SHA-256, replaces only the uniquely resolved `rtk.exe`, verifies the
229
+ new version, and restores and re-verifies the previous bytes on failure. This package-only fallback
230
+ does not widen RFC 0009 or grant permission to mutate agent configuration.
231
+
232
+ RTK has no equivalent machine-readable capability endpoint, so releases newer than the explicitly
233
+ reviewed RTK set remain visible as `unknown-newer` until their consumed contract is checked. See
234
+ [docs/provider-version-compatibility.md](docs/provider-version-compatibility.md).
235
+
236
+ Current experimental candidates include Headroom, mcptoon and GitNexus. Detection or a promising
237
+ benchmark is not enough for promotion. Their campaign assessment is structured evidence for the
238
+ selection gate, not an activation or promotion decision. A candidate must pass structured promotion
239
+ readiness across benchmark capability, category fit, selection evidence, real activation
240
+ verification, managed lifecycle, compatibility/reversibility, project maturity and combined-stack
241
+ validation. Broader context owners also require an explicit admission decision.
242
+
243
+ See [docs/optimizer-priorities.md](docs/optimizer-priorities.md) and
244
+ [RFC 0027](docs/rfcs/0027-optimization-stack-manager.md).
245
+
246
+ ## Stable-stack operating model
247
+
248
+ The intended lifecycle is:
172
249
 
173
- `token-harness` opens the application. `token-harness savings` prints recorded results.
174
- The advanced commands below are implementation tools, not a required user workflow.
250
+ ```text
251
+ discover -> evaluate -> recommend -> install/configure -> verify -> measure
252
+ -> monitor -> update/re-evaluate -> rollback/uninstall
253
+ ```
175
254
 
176
- ## Safety and privacy
255
+ A healthy deterministic component should mostly be left alone. Re-evaluation is useful when an
256
+ agent/optimizer changes version, configuration drift appears, measured value deteriorates, quality
257
+ regresses, workload shape changes materially, or a credible better candidate appears.
177
258
 
178
- Token Harness is conservative by design:
259
+ ## Use Token Harness with an AI agent
179
260
 
180
- - normal read-only commands do not change coding-agent or project configuration;
181
- - `setup --yes`, `apply --yes`, `update --yes`, `rollback --yes`, and
182
- `uninstall --yes` are the explicit CLI configuration-changing forms; the guided UI uses
183
- a reviewed preview and explicit **Approve and apply** instead;
184
- - plans are checked again immediately before they are applied;
185
- - existing files are backed up before a managed write;
186
- - only exact Token Harness-owned entries are removed by `uninstall`;
187
- - newer or untested combinations are reported, not guessed;
188
- - an available provider update outside reviewed compatibility is kept out rather than
189
- forced, and the installed working version stays in place;
190
- - the guided app binds only to 127.0.0.1 and protects its fixed local controls with exact
191
- Host/Origin checks, a per-process anti-forgery token and single-use approval tickets;
192
- - the legacy read-only dashboard and external status seam remain read-only;
193
- - source code, prompts, command contents, credentials, and cookies are not sent to a
194
- Token Harness service.
195
-
196
- Plans, receipts, metrics, and backups stay in the local Token Harness state directory.
197
- See [RFC 0013](docs/rfcs/0013-guided-local-experience.md) for the local browser trust boundary,
198
- [RFC 0004](docs/rfcs/0004-safety-and-installation.md) for the execution model and
199
- [RFC 0006](docs/rfcs/0006-cli-contract.md) for CLI/JSON guarantees.
200
-
201
- ## Supported optimizations
202
-
203
- Token Harness can detect and measure several independent local tools:
204
-
205
- | Provider | Purpose | Management |
206
- | --- | --- | --- |
207
- | [RTK](https://github.com/rtk-ai/rtk) | Shell-command rewriting and output reduction | Managed only for reviewed combinations |
208
- | [HarnessTrim](https://github.com/giuliastro/HarnessTrim) | Deterministic reducers and harness adapters | Managed only for reviewed combinations |
209
- | [cclimits](https://github.com/cruzanstx/cclimits) | Optional live/local quota companion | Read-only; never installed automatically |
210
- | [ccusage](https://github.com/ccusage/ccusage) | Local usage history | Read-only; never installed automatically |
261
+ The browser remains the primary human interface, but the repository also includes a portable Agent
262
+ Skill at [`skills/token-harness/SKILL.md`](skills/token-harness/SKILL.md).
263
+
264
+ If you prefer, you can ask Claude Code or Codex to help with installation and inspection. For
265
+ example:
266
+
267
+ ```text
268
+ Install the latest Token Harness, open it, inspect my coding-agent setup, and explain any proposed
269
+ change before applying it. Do not apply configuration changes without my approval.
270
+ ```
211
271
 
212
- A provider you installed yourself remains yours. Token Harness can adopt observable
213
- configuration without claiming ownership of the executable.
272
+ The skill is deliberately thin: Token Harness remains the deterministic stack/evidence controller.
273
+ The AI does not bypass preview, compatibility checks or explicit approval.
214
274
 
215
- Exact reviewed provider/harness/platform/version combinations are generated in
216
- [docs/matrices.md](docs/matrices.md). A combination outside that table can still be
217
- detected and inspected, but Token Harness will not mutate it.
275
+ The app can also preview enabling that guidance in supported user-level Agent Skills locations.
276
+ Existing custom skill directories are not silently overwritten or adopted. See
277
+ [RFC 0023](docs/rfcs/0023-guided-agent-skill-install.md).
218
278
 
219
- ## Advanced commands
279
+ ## Advanced CLI
220
280
 
221
- Most people do not need this section. Run `token-harness <command> --help` for details.
281
+ Most people do not need these commands. They remain available for automation, debugging and the
282
+ browser controller itself.
222
283
 
223
284
  | Command | Purpose | Changes agent/project config? |
224
285
  | --- | --- | --- |
225
- | `doctor` | Detect harnesses, providers, versions, and problems | No |
286
+ | `doctor` | Detect agents, providers, versions and problems | No |
226
287
  | `budget` | Read authoritative/reported allowance windows | No |
227
- | `context` | Inspect model settings, instructions, and MCP exposure | No |
288
+ | `context` | Inspect model settings, instructions and MCP exposure | No |
228
289
  | `mcp` | Focus on MCP server/tool health | No |
229
290
  | `history` | Summarize local usage through an installed ccusage | No |
230
291
  | `plan` | Prepare exact supported changes | No; stores local plan state |
231
292
  | `apply` | Apply a reviewed stored plan | Yes, only with `--yes` |
232
293
  | `verify` | Check the declared integration tier | No |
233
294
  | `metrics` | Report attributable reducer savings | No |
234
- | `status` | Report pipelines, drift, and importer modes | No |
235
- | `update` | Check/update installed providers; unreviewed targets stay installed | Yes, only with `--yes` |
295
+ | `status` | Report pipelines, drift and importer modes | No |
296
+ | `update` | Check/update reviewed provider packages | Yes, only with `--yes` |
236
297
  | `rollback` | Restore the latest transaction snapshot | Yes, only with `--yes` |
237
298
  | `uninstall` | Remove owned integration entries | Yes, only with `--yes` |
238
299
  | `schedule` | Compare Claude Code and Codex using available evidence | No |
239
- | `handoff` | Build a bounded cross-harness handoff | No |
300
+ | `handoff` | Build a bounded cross-agent handoff | No |
240
301
  | `benchmark*`, `transfer*` | Capture and compare empirical evidence | Local state only |
241
302
 
242
- ## Applying native recommendations
303
+ Need stable machine-readable output? Add `--json`. Need the evidence behind a human summary? Add
304
+ `--verbose`.
243
305
 
244
- `optimize` remains read-only. Review a plan before applying a supported native change:
306
+ The older automation contracts remain available. `ui --json` preserves its existing schema-1
307
+ report; `ui --read-only` opens the legacy read-only UI; `ui --no-open` starts the guided app without
308
+ launching a browser.
245
309
 
246
- ```sh
247
- token-harness plan --harness claude --native-policy --task mechanical --profile economy
248
- token-harness apply --plan <printed-plan-id> --yes
249
- ```
310
+ ### Evaluating an experimental candidate
250
311
 
251
- `apply --plan <id>` restores the reviewed harness/provider selection automatically; you
252
- should not have to repeat `--harness`, `--provider`, `--native-policy`, `--task` or `--profile`.
253
- Run it from the same project as `plan`. Conflicting explicit selectors are rejected, and
254
- actual version, ownership and configuration changes still invalidate the plan. Existing
255
- schema-1 plans remain usable; only their approved actions can execute.
312
+ For normal use, open **Setup -> Experimental tools** and choose **Run standard evaluation**. The app
313
+ keeps a resumable campaign ID for each candidate/harness pair and reads campaign progress, assessment
314
+ and the exact **Next** step directly in the browser. Use **Start baseline capture** or **Start optimized
315
+ capture**, run the requested task in the selected coding agent, then choose **Record outcome** and
316
+ enter the quality/attempt values you actually observed. Normal use no longer requires copying
317
+ `benchmark-start` or `benchmark-finish` commands into a terminal.
256
318
 
257
- The first Claude path supports the **persisted user effort preference** on the reviewed
258
- Claude Code 2.1.261 build. It does not change model, authentication, hooks, endpoint or billing.
259
- `max` is never persisted. Project/local/ancestor settings, custom configuration roots and
260
- known environment/thinking overrides block the change rather than being overwritten. The
261
- preference affects future sessions unless overridden: reopen Claude and check `/effort`.
262
- This is not evidence of a running session's effective effort or a guaranteed quota saving.
319
+ The equivalent advanced CLI flow starts by asking the campaign engine for its current state:
263
320
 
264
- For Codex, the same plan/apply flow manages the existing reviewed reasoning-effort and
265
- verbosity fields through native `config/batchWrite`; project/profile overrides remain yours.
266
- `rollback --yes` restores the complete pre-change files. `uninstall --yes` removes only owned
267
- changes and restores a prior Claude effort preference without undoing unrelated later edits.
321
+ ```sh
322
+ token-harness benchmark-matrix \
323
+ --benchmark-id gitnexus-codex-eval-1 \
324
+ --candidate gitnexus \
325
+ --harness codex
326
+ ```
268
327
 
269
- ## Troubleshooting
328
+ Follow only the **Next** command printed by that report, complete the task honestly, then rerun the
329
+ same `benchmark-matrix` command. Before an optimized run, enable the candidate through its own
330
+ documented workflow. Token Harness records the experiment target but does not treat attribution—or
331
+ the browser acknowledgement—as proof that the candidate was active. For GitNexus, new benchmark
332
+ receipts can additionally verify activation from the harness-native MCP inventory when the GitNexus
333
+ server is observed usable at both task boundaries; missing or ambiguous runtime evidence stays
334
+ unverified.
270
335
 
271
- ### Claude allowance is unavailable
336
+ The selection assessment can become decision-ready after enough evidence across task classes, but it
337
+ still cannot promote a candidate by itself. The remaining lifecycle and combined-stack gates must be
338
+ satisfied separately.
272
339
 
273
- The dashboard now explains whether the optional companion is missing, lacks the safe CLI
274
- flags, cannot find Python, has no usable Claude session, reports an expired session, or returns
275
- an unsupported source. It does not expose credentials, raw companion errors or private paths.
340
+ ### Workload-aware allowance planning
276
341
 
277
- As observed on **September 5, 2026**, npm `cclimits@1.7.0` includes the merged Claude
278
- zero-configuration support and the read-only flags. The latest GitHub Release listing is older
279
- and is not evidence of what npm ships. To check the same path Token Harness uses:
342
+ If you explicitly know the remaining backlog, the advanced CLI can reason about whether that work
343
+ fits the currently observed allowance:
280
344
 
281
345
  ```sh
282
- npm list --global cclimits
283
- cclimits --claude --json --no-cache-write --no-stale-fallback
284
- token-harness budget --harness claude --verbose
346
+ token-harness optimize --harness codex --task standard --tasks-left 5
347
+ token-harness schedule --current codex --candidate claude --task-class standard --tasks-left 5
285
348
  ```
286
349
 
287
- An explicit optional installation/update is `npm install --global cclimits@1.7.0`.
288
- Token Harness does not install it automatically or retry without its read-only flags.
289
- A fresh local Claude cache is shown as **cached**, never promoted to live quota pacing.
290
- A missing observation is not zero remaining allowance. Never paste credentials to debug it.
350
+ For a mixed queued workload:
291
351
 
292
- ### Codex is configured but its hook does not run
352
+ ```sh
353
+ token-harness schedule --current codex --candidate claude \
354
+ --workload mechanical=2,standard=3,hard=1
355
+ ```
293
356
 
294
- `token-harness verify --harness codex --verbose` now reads native `hooks/list` where the
295
- installed app-server exposes it. Disabled, untrusted and modified hooks are distinguished from
296
- an unavailable observation. Trust must still be granted explicitly in Codex. Enabled/trusted
297
- metadata does not prove interception, reduction, or task quality; the integration remains
298
- `config-only` until attributable runtime evidence exists.
357
+ Token Harness does not infer remaining tasks from session length, local tokens or raw provider
358
+ percentages. If the required benchmark/allowance evidence is incomplete, capacity remains unknown.
359
+ See [RFC 0020](docs/rfcs/0020-workload-aware-allowance.md) and
360
+ [RFC 0021](docs/rfcs/0021-mixed-workload-allocation.md).
299
361
 
300
- ### `token-harness` is not found
362
+ ### Applying native recommendations from the CLI
301
363
 
302
- Check that Node is new enough and the package is installed:
364
+ `optimize` remains read-only. The explicit CLI path is review then apply:
303
365
 
304
366
  ```sh
305
- node --version
306
- npm list --global token-harness
367
+ token-harness plan --harness claude --native-policy --task mechanical --profile economy
368
+ token-harness apply --plan <printed-plan-id> --yes
307
369
  ```
308
370
 
309
- Node must be at least 22.13. Reopen the terminal after installation if needed.
371
+ For normal use, prefer the browser workflow.
310
372
 
311
- ### Setup needs attention
373
+ ## Safety and privacy
312
374
 
313
- Run the single command it prints. For technical evidence:
375
+ Token Harness is conservative by design:
376
+
377
+ - opening the app and normal read-only commands do not change agent/project configuration;
378
+ - a browser configuration mutation requires preview and explicit approval;
379
+ - guided candidate capture buttons write only bounded local benchmark state, are CSRF-protected and
380
+ must still match the campaign engine's current step immediately before the write;
381
+ - a browser activation acknowledgement is never treated as verified candidate activation;
382
+ - CLI mutations require their explicit `--yes` form;
383
+ - plans are checked again immediately before apply;
384
+ - existing files are backed up before a managed write;
385
+ - only exact Token Harness-owned entries are removed by uninstall;
386
+ - newer or untested combinations are reported rather than guessed;
387
+ - an available provider update outside reviewed package compatibility is kept out rather than forced;
388
+ - provider package replacement does not bypass the stricter compatibility gate for harness config writes;
389
+ - the guided app binds only to `127.0.0.1` and protects local controls with Host/Origin checks, a
390
+ per-process anti-forgery token and single-use approval tickets;
391
+ - source code, prompts, command contents, credentials and cookies are not sent to a Token Harness
392
+ service.
393
+
394
+ Plans, receipts, metrics and backups stay in the local Token Harness state directory.
395
+
396
+ See [RFC 0013](docs/rfcs/0013-guided-local-experience.md),
397
+ [RFC 0004](docs/rfcs/0004-safety-and-installation.md), and
398
+ [RFC 0006](docs/rfcs/0006-cli-contract.md).
399
+
400
+ ## Updating, checking and undoing
401
+
402
+ Update Token Harness itself:
314
403
 
315
404
  ```sh
316
- token-harness doctor --verbose
405
+ npm install --global token-harness@latest
317
406
  ```
318
407
 
319
- Do not force an unsupported plan. Open an issue with the redacted `--json` result if
320
- you believe the combination should be supported.
408
+ Provider checks and reviewed updates are in **Setup -> Checks and maintenance**. From the advanced
409
+ CLI, first preview and then explicitly apply the provider-package updates:
321
410
 
322
- ### `update` finds a newer version but keeps the installed one
411
+ ```sh
412
+ token-harness update
413
+ token-harness update --yes
414
+ ```
323
415
 
324
- That is normally a safety decision, not a failed installation. Token Harness found a
325
- newer provider release but does not yet have reviewed compatibility evidence for the
326
- active provider × harness × platform combination. Keep using the installed version; no
327
- manual upgrade is required.
416
+ `update` replaces only installed providers whose target is inside the reviewed provider-package
417
+ policy. HarnessTrim uses npm and captures the previous global version for rollback. On native
418
+ Windows RTK prefers WinGet; when WinGet cannot yet reach the reviewed target, Token Harness can use
419
+ the verified official GitHub Windows x64 release fallback described above. That fallback verifies the
420
+ published digest and post-update version and restores the previous executable on failure.
328
421
 
329
- ### Verification says `not-exercised`
422
+ Remove only Token Harness-owned integration entries:
423
+
424
+ ```sh
425
+ token-harness uninstall --yes
426
+ ```
330
427
 
331
- Restart the coding agent, use it for one normal command, and run:
428
+ Restore complete files from the latest committed transaction snapshot:
332
429
 
333
430
  ```sh
334
- token-harness verify
431
+ token-harness rollback --yes
335
432
  ```
336
433
 
337
- No observed operation is different from a failed integration, so Token Harness reports
338
- the two states separately.
434
+ `rollback` is whole-file time travel and can also revert later manual edits to those files. Prefer
435
+ `uninstall` when you only want to remove Token Harness-owned entries.
436
+
437
+ ## Troubleshooting
438
+
439
+ ### Claude allowance is unavailable
339
440
 
340
- ## Updating or undoing
441
+ The app explains whether the optional cclimits companion is missing, too old for safe read-only
442
+ flags, cannot find Python, has no usable Claude session, reports an expired session, or returns an
443
+ unsupported source. It does not expose credentials, raw companion errors or private paths.
341
444
 
342
- Update the CLI:
445
+ For the same technical evidence in the terminal:
343
446
 
344
447
  ```sh
345
- npm install --global token-harness@latest
346
- token-harness setup
448
+ npm list --global cclimits
449
+ cclimits --claude --json --no-cache-write --no-stale-fallback
450
+ token-harness budget --harness claude --verbose
347
451
  ```
348
452
 
349
- Remove only Token Harness-owned integration entries:
453
+ A missing observation is not zero remaining allowance. Never paste credentials to debug it.
454
+
455
+ ### `token-harness` is not found
350
456
 
351
457
  ```sh
352
- token-harness uninstall --yes
458
+ node --version
459
+ npm list --global token-harness
353
460
  ```
354
461
 
355
- Restore complete files from the latest committed transaction snapshot:
462
+ Node must be at least 22.13. Reopen the terminal after installation if needed.
463
+
464
+ ### Setup or verification needs attention
465
+
466
+ Use the action shown in Dashboard or Setup. For technical evidence:
356
467
 
357
468
  ```sh
358
- token-harness rollback --yes
469
+ token-harness doctor --verbose
470
+ token-harness verify --verbose
359
471
  ```
360
472
 
361
- `rollback` is whole-file time travel, so it can also revert later manual edits to those
362
- files. Prefer `uninstall` when you only want to remove Token Harness-owned entries.
473
+ Do not force an unsupported plan. A newer version outside reviewed compatibility is normally a
474
+ safety limitation, not a reason to overwrite the known-working installation.
363
475
 
364
- ## Develop from source
476
+ ## Run the current source
477
+
478
+ From an existing clone:
479
+
480
+ ```sh
481
+ npm start
482
+ ```
483
+
484
+ For a fresh clone:
365
485
 
366
486
  ```sh
367
487
  git clone https://github.com/giuliastro/token-harness.git
368
488
  cd token-harness
369
- corepack enable
370
- pnpm install
371
- pnpm typecheck
372
- pnpm lint
373
- pnpm test
374
- pnpm build
375
- pnpm smoke
376
- pnpm package
377
- pnpm smoke:install
489
+ npx --yes pnpm@10.33.4 install --frozen-lockfile
490
+ npm start
378
491
  ```
379
492
 
380
- Read [PLAN.md](PLAN.md) and the accepted [RFCs](docs/rfcs) before changing public
381
- behavior or architecture.
493
+ This runs the clone. An unmerged branch or unpublished `main` change is not automatically available
494
+ through `token-harness@latest`.
382
495
 
383
- ## License
496
+ ## Development
384
497
 
385
- [Apache License 2.0](LICENSE). Referenced provider tools are independent projects with
386
- their own licenses.
498
+ ```sh
499
+ git clone https://github.com/giuliastro/token-harness.git
500
+ cd token-harness
501
+ npx --yes pnpm@10.33.4 install --frozen-lockfile
502
+ npx --yes pnpm@10.33.4 typecheck
503
+ npx --yes pnpm@10.33.4 lint
504
+ npx --yes pnpm@10.33.4 test
505
+ npx --yes pnpm@10.33.4 build
506
+ npx --yes pnpm@10.33.4 smoke
507
+ npx --yes pnpm@10.33.4 package
508
+ npx --yes pnpm@10.33.4 smoke:install
509
+ ```
387
510
 
388
- ### Loading, impact and sharing
511
+ Using `corepack enable` is optional. On a system-wide Windows Node installation it can require
512
+ administrator permission to modify `C:\Program Files\nodejs`; the `npx pnpm@10.33.4` form above
513
+ does not require that Corepack shim write.
389
514
 
390
- The dashboard shows animated, named checks while it reads your setup. Agent cards and saved
391
- reduction records appear as they are ready; a slow allowance check does not hide the results.
392
- Refreshing keeps previous readings visible until newer ones arrive. Errors and waiting are
393
- explicit, and reduced-motion preferences disable animation without removing status text.
515
+ Before changing public behavior or architecture, read
516
+ [RFC 0027](docs/rfcs/0027-optimization-stack-manager.md),
517
+ [docs/optimizer-priorities.md](docs/optimizer-priorities.md),
518
+ [docs/release-readiness.md](docs/release-readiness.md), [PLAN.md](PLAN.md), and the accepted
519
+ [RFCs](docs/rfcs).
394
520
 
395
- A result can now say **"65% less tool output"**, with its source, before/after values and count
396
- of recorded changed outputs immediately beside it. An estimate says **"Estimated"**. This
397
- percentage describes only those recorded outputs, not your whole coding session, subscription
398
- allowance or money. Provider rows remain separate; negative results and errors remain visible.
521
+ ## License
399
522
 
400
- Choose **Share result** to preview the exact summary and a locally generated image. **Open X
401
- draft** prepares a short post. **Open Reddit** prepares a title/link; copy the summary into a
402
- text post and choose a community yourself. **Copy for Discord** prepares a message to paste
403
- in your chosen channel. **Save image** creates a PNG you can attach yourself. Nothing is
404
- posted or uploaded automatically, and sharing excludes private paths, code, prompts and
405
- account/allowance information. An open share preview stays fixed even if readings update.
523
+ [Apache License 2.0](LICENSE). Referenced provider tools are independent projects with their own
524
+ licenses.