token-harness 0.1.10 → 0.1.12

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (4) hide show
  1. package/README.md +347 -324
  2. package/package.json +1 -1
  3. package/sbom.json +3 -3
  4. package/token-harness.mjs +20165 -9047
package/README.md CHANGED
@@ -1,480 +1,503 @@
1
1
  # Token Harness
2
2
 
3
- **Make Claude Code and Codex easier to understand and use efficiently.**
3
+ **Build, verify and measure an optimization stack for Claude Code and Codex.**
4
4
 
5
- Token Harness checks your coding agents, shows subscription allowance when it can be
6
- observed reliably, reduces avoidable context overhead, and recommends useful actions.
7
- It runs locally and never presents local token estimates as subscription quota.
5
+ Token Harness is a local **optimization stack manager**. It checks your coding agents, manages the
6
+ optimization components it can safely own, keeps evaluation evidence separate from managed lifecycle,
7
+ verifies the result, and reports savings only when it has evidence to support them.
8
8
 
9
- ## Open it. Approve setup. Keep coding.
9
+ It is not another coding agent and it does not replace specialized projects such as RTK or
10
+ HarnessTrim.
10
11
 
11
- You need [Node.js 22.13 or newer](https://nodejs.org/) and an installed, signed-in
12
- Claude Code or Codex. Install Token Harness, then open it:
12
+ ## Start here
13
+
14
+ Requirements:
15
+
16
+ - Node.js 22.13 or newer;
17
+ - Claude Code or Codex installed;
18
+ - the coding agent you want to use already signed in.
19
+
20
+ Install and open Token Harness:
13
21
 
14
22
  ```sh
15
23
  npm install --global token-harness@latest
16
24
  token-harness
17
25
  ```
18
26
 
19
- The browser is now the primary interface. **Set up automatically** checks both agents and
20
- prepares the supported integration changes. It describes each change in plain language.
21
- Choose **Approve and apply** to apply the reviewed configuration with backups and verification.
22
- There are no plan IDs to copy and no daily command sequence to remember.
27
+ That is the normal human workflow. The browser app is the primary interface; there is no daily list
28
+ of CLI commands to memorize.
23
29
 
24
- Already configured? The app shows your existing integrations without replacing them.
25
- An absent provider or an unreviewed version combination is explained rather than installed
26
- or forced silently. Automatic setup covers the reviewed integration paths, not every possible
27
- provider/version. Token Harness does not install Claude Code or Codex or log you in.
30
+ ### First run
28
31
 
29
- ### Daily use
32
+ 1. Open **Dashboard** and let Token Harness inspect the current setup.
33
+ 2. If setup is incomplete, choose **Open setup**.
34
+ 3. In **Setup**, work from top to bottom:
35
+ - Coding agents
36
+ - Managed optimizers
37
+ - Optional agent tuning
38
+ - Checks and maintenance
39
+ 4. For Claude Code or Codex, choose **Review baseline** when the RTK + HarnessTrim setup is available.
40
+ 5. Read the exact proposed changes. **Apply reviewed setup** appears only when there is a concrete
41
+ safe plan to apply.
42
+ 6. Keep using Claude Code or Codex normally.
43
+ 7. Open **Results** when you want to see what Token Harness can actually prove.
30
44
 
31
- Continue launching `claude` or `codex` as usual. Supported output integrations operate in the
32
- agent, not in the dashboard. You can close the page and its terminal without disabling those
33
- integrations. Open `token-harness` whenever you want to see results; the visible dashboard
34
- imports available provider records and refreshes its readings automatically.
45
+ Opening the app does not change your configuration. Read-only checks stay read-only, and a managed
46
+ write requires an explicit review and approval.
35
47
 
36
- **Recorded savings** shows retained history across locally recorded projects, with date bounds,
37
- provider, measurement class, units, changed-output counts, and before/after values. It does not
38
- add incompatible provider figures together. Negative results remain visible. No telemetry is
39
- shown as **not measured**, never a reassuring zero or an invented subscription saving.
40
- Some provider records may predate Token Harness; locally stored records are not guaranteed
41
- complete lifetime history. RTK history is imported directly. HarnessTrim project-local records
42
- must have been imported from their project, or exposed through a configured known metrics path;
43
- the app does not crawl your disk looking for private projects.
48
+ ## The three views
44
49
 
45
- For a terminal-only summary, the one command is:
50
+ ### Dashboard
46
51
 
47
- ```sh
48
- token-harness savings
49
- ```
52
+ Dashboard answers the questions that matter first:
50
53
 
51
- Optional windows are `--since 7d` and `--since 30d`. The advanced `metrics` command remains
52
- project-scoped; opening the app from its installation folder does not change the savings scope.
54
+ - is my setup ready;
55
+ - which coding agents and managed optimizers are active;
56
+ - what should I do next;
57
+ - what value has actually been measured;
58
+ - whether quality or an integration needs attention.
53
59
 
54
- ### Find what you need
60
+ The headline cards deliberately distinguish measured evidence from unknown values. Missing evidence
61
+ is never shown as zero savings.
55
62
 
56
- The app has three tabs: **Overview** for agents and recorded savings, **Rules & settings**
57
- for one agent's rules at a time, and **Activity** for checks and guarded undo. Theme follows
58
- your system; the header also offers light and dark modes.
63
+ ### Setup
59
64
 
60
- Each rule shows what was actually **observed**, separately from how it works. Actions sit
61
- next to the relevant state: **Adjust reasoning** opens the supported review/apply flow;
62
- **Change in Claude/Codex** explains native steps when an automatic write is not available.
63
- Missing allowance data and measurement records have their own setup/help actions.
65
+ Setup is one ordered workflow instead of a collection of unrelated actions.
64
66
 
65
- A saved Claude effort can be displayed even on an unreviewed CLI version. That does not
66
- admit automatic writes on that version. **No saved preference** means the user field is
67
- absent, not that reasoning is disabled. A failed read is a separate state with a cause.
68
- The displayed preference is not a live reading of an already-running session.
67
+ **1. Coding agents**
69
68
 
70
- ### The rules are visible
69
+ Token Harness currently supports guided setup for Claude Code and Codex. Agent details also show
70
+ useful read-only allowance and connected-tool observations when available.
71
71
 
72
- **Rules & settings** explains each configured rule: what it does, why it is used,
73
- its mode, and the evidence available. Automatic integrations, persistent preferences,
74
- observations and features that are not enabled are explicitly distinguished.
72
+ **2. Managed optimizers**
75
73
 
76
- RTK's supported command integration can reduce output automatically. HarnessTrim can use
77
- adapters or skills/instructions, depending on the installation; skills-only is not a transparent
78
- hook, and the agent must actually use the reducer. Configured never means every command was
79
- intercepted. Provider telemetry and its exact/estimated classification remain separate evidence.
74
+ The production baseline remains **RTK + HarnessTrim** on individually reviewed combinations.
75
+ Token Harness also exposes **mcptoon**, **GitNexus** and **Headroom** as optional managed integrations on their exact
76
+ reviewed lifecycle rows; enabling any of them does not make it part of the production baseline or create a
77
+ savings claim. Token Harness tracks the exact combined provider set separately: if no combined-stack
78
+ review is recorded, Setup says so and keeps the stack incomplete rather than inferring compatibility
79
+ from healthy individual checks. Token Harness can prepare their integration transactionally, show the
80
+ exact plan, apply it only after approval, verify it, and remove only configuration it owns.
80
81
 
81
- **Optional: match reasoning to your work** lets you choose the agent and the type of work
82
- without learning CLI flags. It previews the actual supported effort/verbosity change, then
83
- applies only after approval. **This is a persistent preference for future sessions, not an
84
- automatic per-task switch.** It does not switch models, billing, login, or hook trust. The
85
- baseline automatic setup never guesses a task or quietly lowers reasoning.
82
+ For maintainers validating the combined stack, `token-harness stack-review --json` captures the exact
83
+ configured provider versions and managed harness sets and reuses the existing passive `verify`
84
+ evidence for exact provider/harness pairs. Runtime evidence is credited only when it can be
85
+ attributed to that harness: HarnessTrim uses its native event harness field, while provider-wide
86
+ telemetry can be attributed by exclusion only when one harness is wired. With multiple harnesses, an
87
+ unattributable receipt stays **Unavailable** instead of being copied across rows. A `not-exercised`
88
+ result does not promise that ordinary agent use will create a receipt: the provider must record a
89
+ qualifying operation, and for reducers that means a real reduction. `stack-review` does not run an
90
+ active canary or spend a model call, and it never makes the compatibility decision itself. The
91
+ shipped combined-review registry stays empty until the captured configuration has enough real
92
+ verification evidence, has been benchmarked together, and has been deliberately reviewed; see
93
+ `docs/combined-stack-reviews.md`.
86
94
 
87
- **Check integrations** performs the existing integration checks from the UI.
88
- **Undo last change**, available after an application in that dashboard session, previews a
89
- whole-file backup restoration. It refuses to undo a newer unrelated transaction. It restores
90
- only the last successful agent transaction; manual edits to those same files after that
91
- transaction would also be restored, as the confirmation explains.
95
+ **3. Evaluation evidence (maintainers)**
92
96
 
93
- ### Run the current source
97
+ Selection campaigns and historical candidate evidence remain available for maintainers, but they are
98
+ not a separate novice setup product. Once a tool has a reviewed managed lifecycle, Setup presents it in
99
+ the same **Optimization Stack** with its exact state, version, prerequisites, harness scope and safe
100
+ Apply/Remove controls. Evaluation evidence never turns into a savings or compatibility claim by itself.
94
101
 
95
- From an existing clone, after installing its dependencies, one command builds and opens the app:
102
+ Headroom is config-only managed on its exact reviewed 0.37.0 row. Token Harness can own the narrow
103
+ Claude Code or Codex MCP registration after the pinned CLI is present; it deliberately does not
104
+ bootstrap `uv`/Python, start wrappers/proxies, or claim runtime compression savings. GitNexus likewise
105
+ does not auto-index repositories, and its noncommercial license boundary stays visible.
96
106
 
97
- ```sh
98
- npm start
99
- ```
107
+ **4. Optional agent tuning**
100
108
 
101
- For a fresh clone:
109
+ Reasoning preferences are separate from optimizer installation. They are persistent agent settings,
110
+ not hidden per-task switches, and are changed only through the normal preview/apply flow.
102
111
 
103
- ```sh
104
- git clone https://github.com/giuliastro/token-harness.git
105
- cd token-harness
106
- npx --yes pnpm@10.33.4 install --frozen-lockfile
107
- npm start
108
- ```
112
+ **5. Checks and maintenance**
109
113
 
110
- This uses the clone, not an older global installation. An unmerged branch or unpublished main
111
- change is not automatically available through `token-harness@latest`.
114
+ Read-only integration checks, update checks and safe removal/undo controls live here. An update
115
+ outside reviewed compatibility is not forced.
112
116
 
113
- ### Advanced and AI-assisted use
117
+ ### Results
114
118
 
115
- The long CLI flag combinations are **not** the normal human interface. They are the controller API
116
- used by the app, automation, and optionally by a coding harness. Humans can keep using the browser
117
- and the two entry points above.
119
+ Results keeps evidence separate from estimates. Depending on what is actually observable, it can
120
+ show:
118
121
 
119
- For in-session use, this repository now includes a portable Agent Skill at
120
- [`skills/token-harness/SKILL.md`](skills/token-harness/SKILL.md). A compatible Claude Code or Codex
121
- skill mechanism can load it on demand, after which you can simply ask the harness to **use Token
122
- Harness for this task**. The skill classifies substantial work conservatively, calls the existing
123
- local `--json` optimizer at meaningful task boundaries, and can use explicit workload scheduling
124
- when you have actually supplied a backlog. It does not run Token Harness before every tool call.
122
+ - recorded optimizer output reduction;
123
+ - authoritative paired 5-hour / 7-day allowance evidence;
124
+ - API cost only when billed-token evidence and a verified price basis exist;
125
+ - paired quality evidence;
126
+ - experimental candidate evidence and campaign selection assessment when available;
127
+ - recent checks and changes from the current local app session.
125
128
 
126
- The skill is deliberately thin: Token Harness remains the deterministic policy engine. No MCP
127
- server, background model, persistent agent, or second quota formula is added just to make this work.
128
- Local tokens are still not subscription quota, and raw Claude/Codex percentages are still not a
129
- common currency.
129
+ **Not measured** means exactly that. Token Harness does not turn local token estimates into fake
130
+ subscription minutes, money or quota savings.
130
131
 
131
- Agent use is read-only by default. If Token Harness recommends a persistent model, reasoning, or
132
- verbosity change, the harness must first build a reviewed plan, explain the exact proposed mutation,
133
- and wait for your explicit approval before `apply`. A saved preference may affect future sessions;
134
- it is not silently presented as a live change to the current session.
132
+ ## Daily use
135
133
 
136
- The guided app can preview **Enable in-session guidance** for each detected Claude Code or
137
- Codex installation. It installs the same portable skill into the agent's documented user-level
138
- Agent Skills directory through the normal transactional plan/apply path. The agent card then shows
139
- the live guidance state: **Enabled** when the exact skill is currently owned by Token Harness,
140
- **Enabled externally** for a byte-identical user-owned skill, or an explicit not-enabled,
141
- custom/conflict, or unavailable state. Existing `token-harness` skill directories are never
142
- overwritten or silently adopted, and matching bytes alone never create an ownership claim. The
143
- browser remains fully usable without any skill or second AI subscription. See
144
- [RFC 0023](docs/rfcs/0023-guided-agent-skill-install.md) for the install and ownership boundary.
134
+ Keep launching `claude` or `codex` as usual. Deterministic optimizers that are installed, verified
135
+ and still beneficial are intended to remain enabled.
145
136
 
146
- The older automation contracts remain available: `setup`, `optimize`, `plan`, `apply`,
147
- `verify`, `metrics`, `rollback`, and their JSON reports. `ui --json` preserves its existing
148
- schema-1 report; `ui --read-only` opens the legacy read-only dashboard. `ui --no-open` starts
149
- the guided app without launching a browser. Stop either local server with Ctrl+C. See
150
- [RFC 0022](docs/rfcs/0022-agent-native-skill.md) for the agent-facing safety boundary.
137
+ Token Harness does **not** need to stay open, does not need a permanent background daemon, and does
138
+ not need to decide before every command whether an optimizer should run.
151
139
 
152
- ## What normal output looks like
140
+ Open `token-harness` when you want to inspect health/results, review a setup change, check an update,
141
+ verify integrations or re-evaluate the stack after a meaningful version/configuration change.
153
142
 
154
- A healthy final check is intentionally short:
143
+ The app does not periodically reload the whole setup. A full read happens on initial open or when you
144
+ choose **Refresh**. Existing readings remain visible while a refresh runs. Applying a reviewed change
145
+ marks the displayed data as previous state instead of immediately launching another expensive full
146
+ read.
155
147
 
156
- ```text
157
- TOKEN HARNESS - READY
148
+ ## What counts as savings
158
149
 
159
- WHAT WORKS
160
- Codex: configured (0.146.0)
161
- HarnessTrim: active on Codex
150
+ Token Harness keeps different evidence classes separate.
162
151
 
163
- CHANGES
164
- Nothing changed.
152
+ **Recorded output savings** are attributable reducer measurements. Providers, units and measurement
153
+ classes are not silently added together. Negative results and errors remain visible.
165
154
 
166
- NEXT STEP
167
- Use your coding agent normally; configured optimizers run automatically.
168
- ```
155
+ **5h / 7d allowance savings** require authoritative paired before/after allowance evidence. A
156
+ five-hour percentage may also be expressed as the equivalent share of that 300-minute allowance
157
+ window. Weekly quota is not converted into seven days of wall-clock compute.
169
158
 
170
- A newer-than-tested combination is not presented as if the whole setup were broken:
159
+ **API cost** stays **Not measured yet** until attributable billed input/output tokens and a verified
160
+ model-price basis are available.
171
161
 
172
- ```text
173
- TOKEN HARNESS - READY WITH LIMITATIONS
174
-
175
- WHAT WORKS
176
- Claude Code: configured
177
- RTK: active on Claude Code
162
+ **Quality** is measured independently. A measured regression blocks a positive allowance-saving
163
+ claim rather than letting a smaller token number win by itself.
178
164
 
179
- NEXT STEP
180
- token-harness verify
181
- You can keep working; verify the active integrations when convenient.
182
- ```
165
+ Upstream benchmark numbers are useful for deciding what to test; they are never copied directly into
166
+ your savings total.
183
167
 
184
- Need the evidence behind a summary? Add `--verbose`:
168
+ For a terminal-only savings summary:
185
169
 
186
170
  ```sh
187
- token-harness doctor --verbose
171
+ token-harness savings
188
172
  ```
189
173
 
190
- Need stable machine-readable output for automation? Add `--json`:
174
+ Optional windows are `--since 7d` and `--since 30d`.
191
175
 
192
- ```sh
193
- token-harness doctor --json
194
- token-harness ui --json
195
- ```
176
+ ## Current optimization stack
177
+
178
+ Token Harness prefers thin integrations around strong specialized projects instead of copying their
179
+ algorithms into this repository.
180
+
181
+ | Component | Role | Management |
182
+ | --- | --- | --- |
183
+ | [RTK](https://github.com/rtk-ai/rtk) | Shell/tool output reduction | Managed on reviewed combinations |
184
+ | [HarnessTrim](https://github.com/giuliastro/HarnessTrim) | Deterministic output/context reduction | Managed first-party integration |
185
+ | mcptoon | MCP discovery / compact manifest guidance | Optional managed integration on exact reviewed 0.7.10 rows; no savings assumed |
186
+ | GitNexus | Repository graph / MCP context | Optional managed Claude integration for already-installed 1.6.12; license review required |
187
+ | Headroom | Local MCP context compression/retrieval | Optional config-only managed Claude/Codex integration for already-installed 0.37.0; package prerequisite stays user-owned |
188
+ | [cclimits](https://github.com/cruzanstx/cclimits) | Optional Claude allowance evidence | Read-only evidence; not an optimizer |
189
+ | [ccusage](https://github.com/ccusage/ccusage) | Local usage history | Read-only evidence; never subscription quota |
190
+
191
+ Provider compatibility is deliberately **not pinned forever to the first fixture version**. The
192
+ current compatibility policy includes RTK **0.49.0** (source-contract reviewed; the latest live
193
+ Windows harness-mutation fixture is 0.48.0) and HarnessTrim **0.3.0**. Newer HarnessTrim builds can
194
+ be accepted without another hard-coded version bump when their executable version matches their
195
+ machine-readable `capabilities` version and the semantic surface/write-set comparison reports no
196
+ drift.
197
+
198
+ Provider **package updates are separate from harness configuration writes**. `token-harness update`
199
+ can replace a reviewed provider target without requiring an exact historical Claude/Codex fixture
200
+ for that package version; exact compatibility rows still gate any later managed agent-config
201
+ mutation. HarnessTrim updates use its reviewed npm channel and capture the previous global version
202
+ for rollback. On native Windows RTK still prefers WinGet, but when that catalog is behind the
203
+ reviewed 0.49.0 target Token Harness can fall back to the exact official GitHub Windows x64 release:
204
+ it verifies GitHub's published SHA-256, replaces only the uniquely resolved `rtk.exe`, verifies the
205
+ new version, and restores and re-verifies the previous bytes on failure. This package-only fallback
206
+ does not widen RFC 0009 or grant permission to mutate agent configuration.
207
+
208
+ RTK has no equivalent machine-readable capability endpoint, so releases newer than the explicitly
209
+ reviewed RTK set remain visible as `unknown-newer` until their consumed contract is checked. See
210
+ [docs/provider-version-compatibility.md](docs/provider-version-compatibility.md).
211
+
212
+ Historical evaluation evidence remains available for mcptoon, GitNexus and Headroom. Detection or a promising
213
+ benchmark is not enough for a production-stack promotion or savings claim. Their campaign assessment is structured evidence for the
214
+ selection gate, not an activation or promotion decision. A candidate must pass structured promotion
215
+ readiness across benchmark capability, category fit, selection evidence, real activation
216
+ verification, managed lifecycle, compatibility/reversibility, project maturity and combined-stack
217
+ validation. Broader context owners also require an explicit admission decision.
218
+
219
+ See [docs/optimizer-priorities.md](docs/optimizer-priorities.md) and
220
+ [RFC 0027](docs/rfcs/0027-optimization-stack-manager.md).
221
+
222
+ ## Stable-stack operating model
223
+
224
+ The intended lifecycle is:
196
225
 
197
- `--json` keeps the complete schema-1 result and diagnostics; it is not shortened.
226
+ ```text
227
+ discover -> evaluate -> recommend -> install/configure -> verify -> measure
228
+ -> monitor -> update/re-evaluate -> rollback/uninstall
229
+ ```
198
230
 
199
- ## Two entry points to remember
231
+ A healthy deterministic component should mostly be left alone. Re-evaluation is useful when an
232
+ agent/optimizer changes version, configuration drift appears, measured value deteriorates, quality
233
+ regresses, workload shape changes materially, or a credible better candidate appears.
200
234
 
201
- `token-harness` opens the application. `token-harness savings` prints recorded results.
202
- The advanced commands below are implementation tools, not a required user workflow.
235
+ ## Use Token Harness with an AI agent
203
236
 
204
- ## Safety and privacy
237
+ The browser remains the primary human interface, but the repository also includes a portable Agent
238
+ Skill at [`skills/token-harness/SKILL.md`](skills/token-harness/SKILL.md).
205
239
 
206
- Token Harness is conservative by design:
240
+ If you prefer, you can ask Claude Code or Codex to help with installation and inspection. For
241
+ example:
207
242
 
208
- - normal read-only commands do not change coding-agent or project configuration;
209
- - `setup --yes`, `apply --yes`, `update --yes`, `rollback --yes`, and
210
- `uninstall --yes` are the explicit CLI configuration-changing forms; the guided UI uses
211
- a reviewed preview and explicit **Approve and apply** instead;
212
- - plans are checked again immediately before they are applied;
213
- - existing files are backed up before a managed write;
214
- - only exact Token Harness-owned entries are removed by `uninstall`;
215
- - newer or untested combinations are reported, not guessed;
216
- - an available provider update outside reviewed compatibility is kept out rather than
217
- forced, and the installed working version stays in place;
218
- - the guided app binds only to 127.0.0.1 and protects its fixed local controls with exact
219
- Host/Origin checks, a per-process anti-forgery token and single-use approval tickets;
220
- - the legacy read-only dashboard and external status seam remain read-only;
221
- - source code, prompts, command contents, credentials, and cookies are not sent to a
222
- Token Harness service.
223
-
224
- Plans, receipts, metrics, and backups stay in the local Token Harness state directory.
225
- See [RFC 0013](docs/rfcs/0013-guided-local-experience.md) for the local browser trust boundary,
226
- [RFC 0004](docs/rfcs/0004-safety-and-installation.md) for the execution model and
227
- [RFC 0006](docs/rfcs/0006-cli-contract.md) for CLI/JSON guarantees.
228
-
229
- ## Supported optimizations
230
-
231
- Token Harness can detect and measure several independent local tools:
232
-
233
- | Provider | Purpose | Management |
234
- | --- | --- | --- |
235
- | [RTK](https://github.com/rtk-ai/rtk) | Shell-command rewriting and output reduction | Managed only for reviewed combinations |
236
- | [HarnessTrim](https://github.com/giuliastro/HarnessTrim) | Deterministic reducers and harness adapters | Managed only for reviewed combinations |
237
- | [cclimits](https://github.com/cruzanstx/cclimits) | Optional live/local quota companion | Read-only; never installed automatically |
238
- | [ccusage](https://github.com/ccusage/ccusage) | Local usage history | Read-only; never installed automatically |
243
+ ```text
244
+ Install the latest Token Harness, open it, inspect my coding-agent setup, and explain any proposed
245
+ change before applying it. Do not apply configuration changes without my approval.
246
+ ```
239
247
 
240
- A provider you installed yourself remains yours. Token Harness can adopt observable
241
- configuration without claiming ownership of the executable.
248
+ The skill is deliberately thin: Token Harness remains the deterministic stack/evidence controller.
249
+ The AI does not bypass preview, compatibility checks or explicit approval.
242
250
 
243
- Exact reviewed provider/harness/platform/version combinations are generated in
244
- [docs/matrices.md](docs/matrices.md). A combination outside that table can still be
245
- detected and inspected, but Token Harness will not mutate it.
251
+ The app can also preview enabling that guidance in supported user-level Agent Skills locations.
252
+ Existing custom skill directories are not silently overwritten or adopted. See
253
+ [RFC 0023](docs/rfcs/0023-guided-agent-skill-install.md).
246
254
 
247
- ## Advanced commands
255
+ ## Advanced CLI
248
256
 
249
- Most people do not need this section. Run `token-harness <command> --help` for details.
257
+ Most people do not need these commands. They remain available for automation, debugging and the
258
+ browser controller itself.
250
259
 
251
260
  | Command | Purpose | Changes agent/project config? |
252
261
  | --- | --- | --- |
253
- | `doctor` | Detect harnesses, providers, versions, and problems | No |
262
+ | `doctor` | Detect agents, providers, versions and problems | No |
254
263
  | `budget` | Read authoritative/reported allowance windows | No |
255
- | `context` | Inspect model settings, instructions, and MCP exposure | No |
264
+ | `context` | Inspect model settings, instructions and MCP exposure | No |
256
265
  | `mcp` | Focus on MCP server/tool health | No |
257
266
  | `history` | Summarize local usage through an installed ccusage | No |
258
267
  | `plan` | Prepare exact supported changes | No; stores local plan state |
259
268
  | `apply` | Apply a reviewed stored plan | Yes, only with `--yes` |
260
269
  | `verify` | Check the declared integration tier | No |
261
270
  | `metrics` | Report attributable reducer savings | No |
262
- | `status` | Report pipelines, drift, and importer modes | No |
263
- | `update` | Check/update installed providers; unreviewed targets stay installed | Yes, only with `--yes` |
271
+ | `status` | Report pipelines, drift and importer modes | No |
272
+ | `update` | Check/update reviewed provider packages | Yes, only with `--yes` |
264
273
  | `rollback` | Restore the latest transaction snapshot | Yes, only with `--yes` |
265
274
  | `uninstall` | Remove owned integration entries | Yes, only with `--yes` |
266
275
  | `schedule` | Compare Claude Code and Codex using available evidence | No |
267
- | `handoff` | Build a bounded cross-harness handoff | No |
276
+ | `handoff` | Build a bounded cross-agent handoff | No |
268
277
  | `benchmark*`, `transfer*` | Capture and compare empirical evidence | Local state only |
269
278
 
270
- ## Workload-aware allowance planning
279
+ Need stable machine-readable output? Add `--json`. Need the evidence behind a human summary? Add
280
+ `--verbose`.
281
+
282
+ The older automation contracts remain available. `ui --json` preserves its existing schema-1
283
+ report; `ui --read-only` opens the legacy read-only UI; `ui --no-open` starts the guided app without
284
+ launching a browser.
285
+
286
+ ### Evaluation evidence (advanced / maintainers)
271
287
 
272
- If you know how many accepted tasks remain, the advanced CLI can ask whether that backlog fits the
273
- **currently observed** included allowance:
288
+ Evaluation campaigns are an advanced maintainer workflow; managed setup stays in the unified **Optimization Stack**. The app
289
+ keeps a resumable campaign ID for each candidate/harness pair and reads campaign progress, assessment
290
+ and the exact **Next** step directly in the browser. Use **Start baseline capture** or **Start optimized
291
+ capture**, run the requested task in the selected coding agent, then choose **Record outcome** and
292
+ enter the quality/attempt values you actually observed. Normal use no longer requires copying
293
+ `benchmark-start` or `benchmark-finish` commands into a terminal.
294
+
295
+ The equivalent advanced CLI flow starts by asking the campaign engine for its current state. This
296
+ GitNexus example intentionally uses the only currently reviewed campaign row:
274
297
 
275
298
  ```sh
276
- token-harness optimize --harness codex --task standard --tasks-left 5
277
- token-harness schedule --current codex --candidate claude --task-class standard --tasks-left 5
299
+ token-harness benchmark-matrix \
300
+ --benchmark-id gitnexus-claude-eval-1 \
301
+ --candidate gitnexus \
302
+ --harness claude
278
303
  ```
279
304
 
280
- `--tasks-left` is explicit workload intent for a backlog of one task class. Token Harness does not
281
- infer it from `ccusage`, local tokens, session length, or raw provider percentages. A workload-driven
282
- recommendation requires complete project-local benchmark evidence for the exact model + reasoning
283
- effort + verbosity policy in both the five-hour and weekly windows. If that evidence is incomplete,
284
- capacity stays unknown.
305
+ Follow only the **Next** command printed by that report, complete the task honestly, then rerun the
306
+ same `benchmark-matrix` command. Before an optimized run, enable the candidate through its own
307
+ documented workflow. Token Harness records the experiment target but does not treat attribution—or
308
+ the browser acknowledgement—as proof that the candidate was active. For GitNexus on the reviewed
309
+ Claude Code `2.1.269` × GitNexus `1.6.12` × native-Linux row, the harness-native MCP inventory can
310
+ prove only that the GitNexus server was available at both task boundaries. That is not proof Claude
311
+ actually called a GitNexus tool, so the activation-verification promotion gate remains blocked
312
+ without a separate reviewed usage witness. See
313
+ [`docs/candidates/gitnexus-real-campaign.md`](docs/candidates/gitnexus-real-campaign.md).
285
314
 
286
- When evidence proves that the current policy cannot cover the stated backlog, `optimize` protects
287
- capacity instead of spending a quota-derived effort bonus and reports whether the five-hour,
288
- weekly, or both windows are limiting. `schedule` can use the same target to consider the other
289
- harness, but only when that candidate has enough conservative accepted-task capacity and passes the
290
- existing quality, pace, availability, and transfer checks.
315
+ The selection assessment can become decision-ready after enough evidence across task classes, but it
316
+ still cannot promote a candidate by itself. The remaining lifecycle and combined-stack gates must be
317
+ satisfied separately.
291
318
 
292
- ### Mixed task-class backlog
319
+ ### Workload-aware allowance planning
293
320
 
294
- For queued **new tasks** spanning more than one class, give `schedule` the mix explicitly:
321
+ If you explicitly know the remaining backlog, the advanced CLI can reason about whether that work
322
+ fits the currently observed allowance:
295
323
 
296
324
  ```sh
297
- token-harness schedule --current codex --candidate claude \
298
- --workload mechanical=2,standard=3,hard=1
325
+ token-harness optimize --harness codex --task standard --tasks-left 5
326
+ token-harness schedule --current codex --candidate claude --task-class standard --tasks-left 5
299
327
  ```
300
328
 
301
- Mixed mode does not sum per-class task capacities as if they were separate quota buckets. It charges
302
- each proposed task's project-local p75 cost against the same shared five-hour and weekly allowance
303
- of that harness, then returns `stay`, `split`, `switch`, `shortfall`, or
304
- `insufficient-evidence`. Candidate assignments require at least three coherent quality-gated
305
- observations for the exact task class plus complete five-hour and weekly capacity evidence.
306
- Unproven work remains visibly unallocated.
329
+ For a mixed queued workload:
307
330
 
308
- This mode is for queued/new tasks, not an in-progress handoff. Therefore `--workload` is mutually
309
- exclusive with `--task-class`, `--tasks-left`, manual pace/quality flags, and handoff/transfer
310
- flags. The allocator is deterministic and conservative; it does not claim globally optimal routing,
311
- launch either harness, or compare raw Claude and Codex percentages.
331
+ ```sh
332
+ token-harness schedule --current codex --candidate claude \
333
+ --workload mechanical=2,standard=3,hard=1
334
+ ```
312
335
 
313
- No capacity after a future reset is assumed. Re-run the observation after the reset rather than
314
- treating a forecast as provider quota. See [RFC 0020](docs/rfcs/0020-workload-aware-allowance.md)
315
- and [RFC 0021](docs/rfcs/0021-mixed-workload-allocation.md).
336
+ Token Harness does not infer remaining tasks from session length, local tokens or raw provider
337
+ percentages. If the required benchmark/allowance evidence is incomplete, capacity remains unknown.
338
+ See [RFC 0020](docs/rfcs/0020-workload-aware-allowance.md) and
339
+ [RFC 0021](docs/rfcs/0021-mixed-workload-allocation.md).
316
340
 
317
- ## Applying native recommendations
341
+ ### Applying native recommendations from the CLI
318
342
 
319
- `optimize` remains read-only. Review a plan before applying a supported native change:
343
+ `optimize` remains read-only. The explicit CLI path is review then apply:
320
344
 
321
345
  ```sh
322
346
  token-harness plan --harness claude --native-policy --task mechanical --profile economy
323
347
  token-harness apply --plan <printed-plan-id> --yes
324
348
  ```
325
349
 
326
- `apply --plan <id>` restores the reviewed harness/provider selection automatically; you
327
- should not have to repeat `--harness`, `--provider`, `--native-policy`, `--task` or `--profile`.
328
- Run it from the same project as `plan`. Conflicting explicit selectors are rejected, and
329
- actual version, ownership and configuration changes still invalidate the plan. Existing
330
- schema-1 plans remain usable; only their approved actions can execute.
350
+ For normal use, prefer the browser workflow.
331
351
 
332
- The first Claude path supports the **persisted user effort preference** on the reviewed
333
- Claude Code 2.1.261 build. It does not change model, authentication, hooks, endpoint or billing.
334
- `max` is never persisted. Project/local/ancestor settings, custom configuration roots and
335
- known environment/thinking overrides block the change rather than being overwritten. The
336
- preference affects future sessions unless overridden: reopen Claude and check `/effort`.
337
- This is not evidence of a running session's effective effort or a guaranteed quota saving.
352
+ ## Safety and privacy
338
353
 
339
- For Codex, the same plan/apply flow manages the existing reviewed reasoning-effort and
340
- verbosity fields through native `config/batchWrite`; project/profile overrides remain yours.
341
- `rollback --yes` restores the complete pre-change files. `uninstall --yes` removes only owned
342
- changes and restores a prior Claude effort preference without undoing unrelated later edits.
354
+ Token Harness is conservative by design:
343
355
 
344
- ## Troubleshooting
356
+ - opening the app and normal read-only commands do not change agent/project configuration;
357
+ - a browser configuration mutation requires preview and explicit approval;
358
+ - guided candidate capture buttons write only bounded local benchmark state, are CSRF-protected and
359
+ must still match the campaign engine's current step immediately before the write;
360
+ - a browser activation acknowledgement is never treated as verified candidate activation;
361
+ - CLI mutations require their explicit `--yes` form;
362
+ - plans are checked again immediately before apply;
363
+ - existing files are backed up before a managed write;
364
+ - only exact Token Harness-owned entries are removed by uninstall;
365
+ - newer or untested combinations are reported rather than guessed;
366
+ - an available provider update outside reviewed package compatibility is kept out rather than forced;
367
+ - provider package replacement does not bypass the stricter compatibility gate for harness config writes;
368
+ - the guided app binds only to `127.0.0.1` and protects local controls with Host/Origin checks, a
369
+ per-process anti-forgery token and single-use approval tickets;
370
+ - source code, prompts, command contents, credentials and cookies are not sent to a Token Harness
371
+ service.
345
372
 
346
- ### Claude allowance is unavailable
373
+ Plans, receipts, metrics and backups stay in the local Token Harness state directory.
347
374
 
348
- The dashboard now explains whether the optional companion is missing, lacks the safe CLI
349
- flags, cannot find Python, has no usable Claude session, reports an expired session, or returns
350
- an unsupported source. It does not expose credentials, raw companion errors or private paths.
375
+ See [RFC 0013](docs/rfcs/0013-guided-local-experience.md),
376
+ [RFC 0004](docs/rfcs/0004-safety-and-installation.md), and
377
+ [RFC 0006](docs/rfcs/0006-cli-contract.md).
351
378
 
352
- As observed on **September 5, 2026**, npm `cclimits@1.7.0` includes the merged Claude
353
- zero-configuration support and the read-only flags. The latest GitHub Release listing is older
354
- and is not evidence of what npm ships. To check the same path Token Harness uses:
379
+ ## Updating, checking and undoing
380
+
381
+ Update Token Harness itself:
355
382
 
356
383
  ```sh
357
- npm list --global cclimits
358
- cclimits --claude --json --no-cache-write --no-stale-fallback
359
- token-harness budget --harness claude --verbose
384
+ npm install --global token-harness@latest
360
385
  ```
361
386
 
362
- An explicit optional installation/update is `npm install --global cclimits@1.7.0`.
363
- Token Harness does not install it automatically or retry without its read-only flags.
364
- A fresh local Claude cache is shown as **cached**, never promoted to live quota pacing.
365
- A missing observation is not zero remaining allowance. Never paste credentials to debug it.
366
-
367
- ### Codex is configured but its hook does not run
387
+ Provider checks and reviewed updates are in **Setup -> Checks and maintenance**. From the advanced
388
+ CLI, first preview and then explicitly apply the provider-package updates:
368
389
 
369
- `token-harness verify --harness codex --verbose` now reads native `hooks/list` where the
370
- installed app-server exposes it. Disabled, untrusted and modified hooks are distinguished from
371
- an unavailable observation. Trust must still be granted explicitly in Codex. Enabled/trusted
372
- metadata does not prove interception, reduction, or task quality; the integration remains
373
- `config-only` until attributable runtime evidence exists.
390
+ ```sh
391
+ token-harness update
392
+ token-harness update --yes
393
+ ```
374
394
 
375
- ### `token-harness` is not found
395
+ `update` replaces only installed providers whose target is inside the reviewed provider-package
396
+ policy. HarnessTrim uses npm and captures the previous global version for rollback. On native
397
+ Windows RTK prefers WinGet; when WinGet cannot yet reach the reviewed target, Token Harness can use
398
+ the verified official GitHub Windows x64 release fallback described above. That fallback verifies the
399
+ published digest and post-update version and restores the previous executable on failure.
376
400
 
377
- Check that Node is new enough and the package is installed:
401
+ Remove only Token Harness-owned integration entries:
378
402
 
379
403
  ```sh
380
- node --version
381
- npm list --global token-harness
404
+ token-harness uninstall --yes
382
405
  ```
383
406
 
384
- Node must be at least 22.13. Reopen the terminal after installation if needed.
385
-
386
- ### Setup needs attention
387
-
388
- Run the single command it prints. For technical evidence:
407
+ Restore complete files from the latest committed transaction snapshot:
389
408
 
390
409
  ```sh
391
- token-harness doctor --verbose
410
+ token-harness rollback --yes
392
411
  ```
393
412
 
394
- Do not force an unsupported plan. Open an issue with the redacted `--json` result if
395
- you believe the combination should be supported.
413
+ `rollback` is whole-file time travel and can also revert later manual edits to those files. Prefer
414
+ `uninstall` when you only want to remove Token Harness-owned entries.
396
415
 
397
- ### `update` finds a newer version but keeps the installed one
416
+ ## Troubleshooting
398
417
 
399
- That is normally a safety decision, not a failed installation. Token Harness found a
400
- newer provider release but does not yet have reviewed compatibility evidence for the
401
- active provider × harness × platform combination. Keep using the installed version; no
402
- manual upgrade is required.
418
+ ### Claude allowance is unavailable
403
419
 
404
- ### Verification says `not-exercised`
420
+ The app explains whether the optional cclimits companion is missing, too old for safe read-only
421
+ flags, cannot find Python, has no usable Claude session, reports an expired session, or returns an
422
+ unsupported source. It does not expose credentials, raw companion errors or private paths.
405
423
 
406
- Restart the coding agent, use it for one normal command, and run:
424
+ For the same technical evidence in the terminal:
407
425
 
408
426
  ```sh
409
- token-harness verify
427
+ npm list --global cclimits
428
+ cclimits --claude --json --no-cache-write --no-stale-fallback
429
+ token-harness budget --harness claude --verbose
410
430
  ```
411
431
 
412
- No observed operation is different from a failed integration, so Token Harness reports
413
- the two states separately.
414
-
415
- ## Updating or undoing
432
+ A missing observation is not zero remaining allowance. Never paste credentials to debug it.
416
433
 
417
- Update the CLI:
434
+ ### `token-harness` is not found
418
435
 
419
436
  ```sh
420
- npm install --global token-harness@latest
421
- token-harness setup
437
+ node --version
438
+ npm list --global token-harness
422
439
  ```
423
440
 
424
- Remove only Token Harness-owned integration entries:
441
+ Node must be at least 22.13. Reopen the terminal after installation if needed.
442
+
443
+ ### Setup or verification needs attention
444
+
445
+ Use the action shown in Dashboard or Setup. For technical evidence:
425
446
 
426
447
  ```sh
427
- token-harness uninstall --yes
448
+ token-harness doctor --verbose
449
+ token-harness verify --verbose
428
450
  ```
429
451
 
430
- Restore complete files from the latest committed transaction snapshot:
452
+ Do not force an unsupported plan. A newer version outside reviewed compatibility is normally a
453
+ safety limitation, not a reason to overwrite the known-working installation.
454
+
455
+ ## Run the current source
456
+
457
+ From an existing clone:
431
458
 
432
459
  ```sh
433
- token-harness rollback --yes
460
+ npm start
434
461
  ```
435
462
 
436
- `rollback` is whole-file time travel, so it can also revert later manual edits to those
437
- files. Prefer `uninstall` when you only want to remove Token Harness-owned entries.
438
-
439
- ## Develop from source
463
+ For a fresh clone:
440
464
 
441
465
  ```sh
442
466
  git clone https://github.com/giuliastro/token-harness.git
443
467
  cd token-harness
444
- corepack enable
445
- pnpm install
446
- pnpm typecheck
447
- pnpm lint
448
- pnpm test
449
- pnpm build
450
- pnpm smoke
451
- pnpm package
452
- pnpm smoke:install
468
+ npx --yes pnpm@10.33.4 install --frozen-lockfile
469
+ npm start
453
470
  ```
454
471
 
455
- Read [PLAN.md](PLAN.md) and the accepted [RFCs](docs/rfcs) before changing public
456
- behavior or architecture.
472
+ This runs the clone. An unmerged branch or unpublished `main` change is not automatically available
473
+ through `token-harness@latest`.
457
474
 
458
- ## License
475
+ ## Development
459
476
 
460
- [Apache License 2.0](LICENSE). Referenced provider tools are independent projects with
461
- their own licenses.
477
+ ```sh
478
+ git clone https://github.com/giuliastro/token-harness.git
479
+ cd token-harness
480
+ npx --yes pnpm@10.33.4 install --frozen-lockfile
481
+ npx --yes pnpm@10.33.4 typecheck
482
+ npx --yes pnpm@10.33.4 lint
483
+ npx --yes pnpm@10.33.4 test
484
+ npx --yes pnpm@10.33.4 build
485
+ npx --yes pnpm@10.33.4 smoke
486
+ npx --yes pnpm@10.33.4 package
487
+ npx --yes pnpm@10.33.4 smoke:install
488
+ ```
462
489
 
463
- ### Loading, impact and sharing
490
+ Using `corepack enable` is optional. On a system-wide Windows Node installation it can require
491
+ administrator permission to modify `C:\Program Files\nodejs`; the `npx pnpm@10.33.4` form above
492
+ does not require that Corepack shim write.
464
493
 
465
- The dashboard shows animated, named checks while it reads your setup. Agent cards and saved
466
- reduction records appear as they are ready; a slow allowance check does not hide the results.
467
- Refreshing keeps previous readings visible until newer ones arrive. Errors and waiting are
468
- explicit, and reduced-motion preferences disable animation without removing status text.
494
+ Before changing public behavior or architecture, read
495
+ [RFC 0027](docs/rfcs/0027-optimization-stack-manager.md),
496
+ [docs/optimizer-priorities.md](docs/optimizer-priorities.md),
497
+ [docs/release-readiness.md](docs/release-readiness.md), [PLAN.md](PLAN.md), and the accepted
498
+ [RFCs](docs/rfcs).
469
499
 
470
- A result can now say **"65% less tool output"**, with its source, before/after values and count
471
- of recorded changed outputs immediately beside it. An estimate says **"Estimated"**. This
472
- percentage describes only those recorded outputs, not your whole coding session, subscription
473
- allowance or money. Provider rows remain separate; negative results and errors remain visible.
500
+ ## License
474
501
 
475
- Choose **Share result** to preview the exact summary and a locally generated image. **Open X
476
- draft** prepares a short post. **Open Reddit** prepares a title/link; copy the summary into a
477
- text post and choose a community yourself. **Copy for Discord** prepares a message to paste
478
- in your chosen channel. **Save image** creates a PNG you can attach yourself. Nothing is
479
- posted or uploaded automatically, and sharing excludes private paths, code, prompts and
480
- account/allowance information. An open share preview stays fixed even if readings update.
502
+ [Apache License 2.0](LICENSE). Referenced provider tools are independent projects with their own
503
+ licenses.