token-harness 0.1.6 → 0.1.8
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +200 -714
- package/package.json +1 -1
- package/sbom.json +3 -3
- package/token-harness.mjs +1593 -483
package/README.md
CHANGED
|
@@ -1,835 +1,326 @@
|
|
|
1
1
|
# Token Harness
|
|
2
2
|
|
|
3
|
-
|
|
4
|
-
and Codex usage limits, without hiding quality regressions or pretending that opaque subscription
|
|
5
|
-
quota is exactly equivalent to a token count**.
|
|
3
|
+
**Make Claude Code and Codex easier to understand and use efficiently.**
|
|
6
4
|
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
evolving from a token-reduction control plane into a **quota-aware efficiency layer** for coding
|
|
11
|
-
harnesses.
|
|
5
|
+
Token Harness checks your coding agents, shows subscription allowance when it can be
|
|
6
|
+
observed reliably, reduces avoidable context overhead, and recommends useful actions.
|
|
7
|
+
It runs locally and never presents local token estimates as subscription quota.
|
|
12
8
|
|
|
13
|
-
The
|
|
9
|
+
## The easy path
|
|
14
10
|
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
budget;
|
|
18
|
-
2. **context sent to the model** — keep instructions, MCP schemas, conversation history, repository
|
|
19
|
-
context, and tool output as small as possible without removing information needed to finish the
|
|
20
|
-
task correctly.
|
|
11
|
+
You need [Node.js 22.13 or newer](https://nodejs.org/) and at least one signed-in coding
|
|
12
|
+
agent such as Claude Code or Codex.
|
|
21
13
|
|
|
22
|
-
|
|
23
|
-
than the product definition. Native harness controls come first because choosing the right model,
|
|
24
|
-
effort, context shape, and enabled tools can avoid an expensive turn entirely.
|
|
25
|
-
|
|
26
|
-
## What Token Harness should optimize
|
|
27
|
-
|
|
28
|
-
| Layer | Target behavior |
|
|
29
|
-
| --- | --- |
|
|
30
|
-
| Usage-window observability | Read five-hour, weekly, model-specific, or credit-backed limits only from surfaces the harness can actually prove; show reset time, headroom, and burn rate |
|
|
31
|
-
| Native model policy | Prefer economical models/effort for routine work, escalate only for tasks whose expected quality benefit justifies the extra quota |
|
|
32
|
-
| Session hygiene | Detect task boundaries, oversized conversations, and stale context; recommend clear/compact/new-session actions before history dominates each turn |
|
|
33
|
-
| Instruction budget | Keep `CLAUDE.md` / `AGENTS.md` concise and hierarchical instead of injecting one large global instruction file everywhere |
|
|
34
|
-
| MCP/tool budget | Disable irrelevant MCP servers, defer schemas where the harness supports it, and expose only the tools required by the current task |
|
|
35
|
-
| Tool-output budget | Reduce logs, diffs, test output, and repeated command results before they re-enter model context |
|
|
36
|
-
| Cross-harness scheduling | When Claude Code and Codex have independent headroom, recommend the harness that can do the task with the best expected quality per remaining quota |
|
|
37
|
-
| Measurement | Correlate quota delta, model/effort, context size, tool traffic, and task outcome; never convert local token savings into an unproven subscription-quota claim |
|
|
38
|
-
|
|
39
|
-
## Optimization ecosystem
|
|
40
|
-
|
|
41
|
-
The priority order below is deliberately different from the original Token Harness roadmap. A tool
|
|
42
|
-
is high priority only when it helps **included Claude Code/Codex capacity**, not merely API cost,
|
|
43
|
-
provider routing, or benchmark token counts.
|
|
44
|
-
|
|
45
|
-
| Tool or surface | Optimization layer | Token Harness direction |
|
|
46
|
-
| --- | --- | --- |
|
|
47
|
-
| **Claude Code native controls** | Usage status, model/effort choice, context inspection, clear/compact lifecycle | **P0 — build into core**. Prefer native surfaces over third-party credential scraping; discover available models dynamically |
|
|
48
|
-
| **Codex native app-server + profiles** | Structured rate-limit telemetry; model, reasoning effort, verbosity, tool-output and MCP/context controls | **P0 — build into core**. `account/rateLimits/read` is the preferred live quota source when available |
|
|
49
|
-
| [cclimits](https://github.com/cruzanstx/cclimits) | Live/local quota companion across coding tools | **P0 — optional observational companion for Claude**. Token Harness uses only cacheless Claude JSON mode, never reads its credentials, and labels the result `reported` or `cached`. Codex keeps its native app-server reader |
|
|
50
|
-
| [ccusage](https://github.com/ccusage/ccusage) | Local historical token/session/cost analytics for Claude Code and Codex | **P0 — high-priority read-only companion**. Useful for history and baselines; not a substitute for live subscription-limit telemetry |
|
|
51
|
-
| [RTK](https://github.com/rtk-ai/rtk) | Shell-command rewriting and command-output reduction | **P1 — active, keep**. Reduces context that would otherwise be resent on later turns |
|
|
52
|
-
| [HarnessTrim](https://github.com/giuliastro/HarnessTrim) | Deterministic reducers, harness adapters, skills, pipes, and MCP reduction | **P1 — active, keep**. Continue only on proven non-overlapping surfaces |
|
|
53
|
-
| [Lazy MCP](https://gitlab.com/gitlab-org/ai/lazy-mcp) | Load MCP tool schemas only when needed | **P1 — high priority**. Directly attacks schema/context overhead; benchmark against native Codex tool deferral before installing another owner |
|
|
54
|
-
| [Context Mode](https://github.com/mksglu/context-mode) | Keep raw tool/MCP results outside model context | **P2 — alternative broad context owner**. Evaluate against Headroom, not alongside overlapping reducers by default |
|
|
55
|
-
| [Headroom](https://github.com/headroomlabs-ai/headroom) | Compress tool, MCP, file, and RAG payloads | **P2 — alternative broad context owner**. Admit only after pair-specific quality and attribution fixtures |
|
|
56
|
-
| [Dejavu](https://github.com/Salnika/dejavu) | Emit only the delta when command output repeats | **P2 — useful for test/rerun loops** after the normal output-reduction path is measured |
|
|
57
|
-
| [repowise](https://github.com/repowise-dev/repowise) | Retrieve task-specific repository context | **P2 — conditional**. Its MCP overhead must be lower than the repository context it avoids |
|
|
58
|
-
| LiteLLM, Claude Code Router, RouteLLM, LLMRouter, vLLM Semantic Router | Route requests across models/providers | **P3 — overflow/API routing only**. Useful after included quota is exhausted or for explicit external-provider policy; not a core subscription-quota optimizer |
|
|
59
|
-
| [LLMLingua](https://github.com/microsoft/LLMLingua) | Generic prompt compression | **Research only** until a harness-aware lifecycle proves must-keep recall and quality |
|
|
60
|
-
| [Caveman](https://github.com/JuliusBrussee/caveman) | Reduce visible model-output verbosity | **De-prioritized**. Use native verbosity/effort controls first and measure their effect before adding another instruction owner |
|
|
61
|
-
|
|
62
|
-
Candidate status still means that Token Harness neither installs nor configures the tool until its
|
|
63
|
-
installation, conflicts, rollback behavior, verification, quality gates, and metrics attribution are
|
|
64
|
-
implemented and tested. The intake evidence lives in
|
|
65
|
-
[docs/provider-landscape.md](docs/provider-landscape.md).
|
|
66
|
-
|
|
67
|
-
### Why generic routers moved down
|
|
68
|
-
|
|
69
|
-
The original roadmap treated model routers as high priority. For the new objective that is backwards:
|
|
70
|
-
a router can reduce API cost by sending work to another provider while consuming **none of the
|
|
71
|
-
included Claude Code or Codex allowance**, or it can bypass subscription authentication entirely.
|
|
72
|
-
That can be valuable as explicit overflow, but it does not demonstrate better use of the allowance
|
|
73
|
-
the user already paid for.
|
|
74
|
-
|
|
75
|
-
Token Harness should first exploit the harness-native choices that are actually inside the plan:
|
|
76
|
-
model tier, reasoning effort, verbosity, session lifecycle, instruction size, MCP exposure, and tool
|
|
77
|
-
output. External routing becomes an opt-in second budget, never an invisible shortcut.
|
|
78
|
-
|
|
79
|
-
## Quick start
|
|
80
|
-
|
|
81
|
-
Install the CLI:
|
|
82
|
-
|
|
83
|
-
```sh
|
|
84
|
-
npm install --global token-harness
|
|
85
|
-
token-harness --version
|
|
86
|
-
```
|
|
87
|
-
|
|
88
|
-
Then run the complete workflow from the project in which you use your coding agent:
|
|
14
|
+
Open a terminal and paste:
|
|
89
15
|
|
|
90
16
|
```sh
|
|
91
|
-
|
|
92
|
-
token-harness
|
|
93
|
-
|
|
94
|
-
# 2. Inspect subscription headroom and avoidable context. Still read-only.
|
|
95
|
-
token-harness budget
|
|
96
|
-
token-harness context
|
|
97
|
-
token-harness history --since 7d
|
|
98
|
-
token-harness optimize
|
|
99
|
-
|
|
100
|
-
# Optional: Claude live quota can use an already-installed cclimits build that
|
|
101
|
-
# supports --no-cache-write. Token Harness never installs it automatically.
|
|
102
|
-
|
|
103
|
-
# 3. Preview every proposed configuration change.
|
|
104
|
-
token-harness plan
|
|
105
|
-
|
|
106
|
-
# 4. Apply the reviewed plan. This is the first configuration-changing step.
|
|
107
|
-
token-harness apply --yes
|
|
108
|
-
|
|
109
|
-
# 5. Restart the coding agent, then run a normal shell command through it.
|
|
110
|
-
|
|
111
|
-
# 6. Check configuration, real interception evidence, and savings.
|
|
112
|
-
token-harness status
|
|
113
|
-
token-harness verify
|
|
114
|
-
token-harness metrics --since 7d
|
|
17
|
+
npm install --global token-harness@latest
|
|
18
|
+
token-harness setup
|
|
115
19
|
```
|
|
116
20
|
|
|
117
|
-
|
|
118
|
-
there.
|
|
21
|
+
That is the whole first-time check. `setup` tells you:
|
|
119
22
|
|
|
120
|
-
|
|
23
|
+
- which coding agents it found;
|
|
24
|
+
- which optimizers are already active;
|
|
25
|
+
- whether the current setup works;
|
|
26
|
+
- what it changed, if anything;
|
|
27
|
+
- exactly one next step.
|
|
121
28
|
|
|
122
|
-
|
|
29
|
+
The first run does not change Claude Code, Codex, or project configuration. If Token
|
|
30
|
+
Harness finds a supported improvement, it shows the safe plan and may suggest:
|
|
123
31
|
|
|
124
32
|
```sh
|
|
125
|
-
|
|
126
|
-
token-harness budget --harness codex
|
|
127
|
-
|
|
128
|
-
# Read-only: inspect the effective model, reasoning effort, project instructions,
|
|
129
|
-
# MCP servers, and visible tool inventory.
|
|
130
|
-
token-harness context --harness codex
|
|
131
|
-
|
|
132
|
-
# Read-only: combine quota pacing and context pressure into task-specific advice.
|
|
133
|
-
token-harness optimize --harness codex --task standard --profile balanced
|
|
134
|
-
|
|
135
|
-
# Dry run: preview supported Codex-native policy changes.
|
|
136
|
-
token-harness plan --harness codex --native-policy
|
|
33
|
+
token-harness setup --yes
|
|
137
34
|
```
|
|
138
35
|
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
also include reviewed provider actions such as a HarnessTrim skill install.
|
|
36
|
+
`--yes` is always explicit. The change is backed up, applied transactionally, and
|
|
37
|
+
verified. Unsupported combinations are left untouched.
|
|
142
38
|
|
|
143
|
-
|
|
39
|
+
## Open the dashboard
|
|
144
40
|
|
|
145
|
-
|
|
146
|
-
token-harness apply --plan <plan-id> --yes
|
|
147
|
-
token-harness context --harness codex
|
|
148
|
-
token-harness verify --harness codex
|
|
149
|
-
```
|
|
150
|
-
|
|
151
|
-
On supported recent Codex builds, Token Harness uses Codex's native app-server for authoritative
|
|
152
|
-
rate-limit windows, effective configuration, model catalog, MCP inventory, and reviewed native
|
|
153
|
-
configuration writes. It does not infer quota from local token counts.
|
|
154
|
-
|
|
155
|
-
Useful variants:
|
|
41
|
+
After setup, run:
|
|
156
42
|
|
|
157
43
|
```sh
|
|
158
|
-
|
|
159
|
-
token-harness optimize --harness codex --task mechanical --profile economy
|
|
160
|
-
|
|
161
|
-
# Preserve more quality headroom for difficult work.
|
|
162
|
-
token-harness optimize --harness codex --task hard --profile quality
|
|
163
|
-
|
|
164
|
-
# Inspect MCP exposure directly.
|
|
165
|
-
token-harness mcp --harness codex
|
|
166
|
-
|
|
167
|
-
# Compare recent local history when available.
|
|
168
|
-
token-harness history --harness codex --since 7d
|
|
44
|
+
token-harness ui
|
|
169
45
|
```
|
|
170
46
|
|
|
171
|
-
|
|
47
|
+
The dashboard opens in your browser and answers three questions in this order:
|
|
172
48
|
|
|
173
|
-
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
from attributable project-local benchmark evidence. Missing or conflicting evidence returns
|
|
177
|
-
`insufficient-evidence` instead of guessing.
|
|
49
|
+
1. **Can I work normally right now?**
|
|
50
|
+
2. **What is actually active and useful?**
|
|
51
|
+
3. **Is there one action worth taking?**
|
|
178
52
|
|
|
179
|
-
|
|
53
|
+
It shows active/relevant coding agents first, their optimization providers, observable
|
|
54
|
+
allowance windows, and task guidance. Tools that are absent do not get large cards;
|
|
55
|
+
secondary detected tools are kept out of the main path.
|
|
180
56
|
|
|
181
|
-
|
|
182
|
-
token-harness schedule \
|
|
183
|
-
--current claude \
|
|
184
|
-
--candidate codex \
|
|
185
|
-
--task-class hard \
|
|
186
|
-
--handoff-bytes 900
|
|
187
|
-
```
|
|
57
|
+
The page is served only on `127.0.0.1`: no Electron app, account, or cloud service.
|
|
188
58
|
|
|
189
|
-
|
|
59
|
+
### What happens after the dashboard?
|
|
190
60
|
|
|
191
|
-
|
|
192
|
-
token-harness handoff \
|
|
193
|
-
--objective "Finish the scheduler change without weakening quota evidence" \
|
|
194
|
-
--decision "Keep Claude and Codex quota observations provider-local" \
|
|
195
|
-
--changed-file apps/cli/src/schedule-main.ts \
|
|
196
|
-
--validation "pnpm test passes" \
|
|
197
|
-
--unresolved "Need one empirical Codex comparison" \
|
|
198
|
-
--next-action "Run the candidate benchmark in Codex" \
|
|
199
|
-
--max-bytes 2048 > handoff.md
|
|
200
|
-
```
|
|
61
|
+
Usually: **nothing else. You are done.**
|
|
201
62
|
|
|
202
|
-
|
|
203
|
-
|
|
204
|
-
but different harnesses:
|
|
63
|
+
Token Harness is not a launcher and it does not need to stay between you and your coding
|
|
64
|
+
agent. Continue exactly as you normally would:
|
|
205
65
|
|
|
206
66
|
```sh
|
|
207
|
-
|
|
208
|
-
|
|
209
|
-
|
|
210
|
-
--variant baseline \
|
|
211
|
-
--task hard \
|
|
212
|
-
--harness claude
|
|
213
|
-
|
|
214
|
-
# Run the baseline task in Claude Code, then record its observed outcome.
|
|
215
|
-
token-harness benchmark-finish \
|
|
216
|
-
--benchmark-id scheduler-hard-01 \
|
|
217
|
-
--variant baseline \
|
|
218
|
-
--quality passed \
|
|
219
|
-
--attempts 1 \
|
|
220
|
-
--failed-attempts 0
|
|
221
|
-
|
|
222
|
-
# 2. Before the candidate run, start the optimized variant.
|
|
223
|
-
token-harness benchmark-start \
|
|
224
|
-
--benchmark-id scheduler-hard-01 \
|
|
225
|
-
--variant optimized \
|
|
226
|
-
--task hard \
|
|
227
|
-
--harness codex
|
|
228
|
-
|
|
229
|
-
# Give handoff.md to Codex, run the equivalent task, then close the capture.
|
|
230
|
-
token-harness benchmark-finish \
|
|
231
|
-
--benchmark-id scheduler-hard-01 \
|
|
232
|
-
--variant optimized \
|
|
233
|
-
--quality passed \
|
|
234
|
-
--attempts 1 \
|
|
235
|
-
--failed-attempts 0
|
|
236
|
-
|
|
237
|
-
# 3. Evaluate the exact handoff used by the candidate run.
|
|
238
|
-
token-harness transfer \
|
|
239
|
-
--benchmark-id scheduler-hard-01 \
|
|
240
|
-
--handoff-file handoff.md
|
|
241
|
-
|
|
242
|
-
# 4. Persist the immutable transfer verdict for future schedule calls.
|
|
243
|
-
token-harness transfer-record \
|
|
244
|
-
--benchmark-id scheduler-hard-01 \
|
|
245
|
-
--handoff-file handoff.md
|
|
67
|
+
claude
|
|
68
|
+
codex
|
|
69
|
+
opencode
|
|
246
70
|
```
|
|
247
71
|
|
|
248
|
-
|
|
249
|
-
|
|
250
|
-
|
|
251
|
-
positive/non-positive conflict keeps the recommendation unresolved. Historical receipt sizes never
|
|
252
|
-
replace the current `--handoff-bytes` value.
|
|
72
|
+
Use whichever of those you already use. RTK, HarnessTrim, or another configured provider
|
|
73
|
+
runs automatically through that coding agent's integration. You do **not** need a special
|
|
74
|
+
`token-harness run` command.
|
|
253
75
|
|
|
254
|
-
|
|
255
|
-
|
|
256
|
-
`token-harness schedule --help` for the full list.
|
|
76
|
+
You can close the dashboard whenever you want. Use `Ctrl+C` in the terminal to stop its
|
|
77
|
+
local web server. Closing it does not disable configured optimizers.
|
|
257
78
|
|
|
258
|
-
|
|
79
|
+
Open it again later with `token-harness ui` when you want a status check. If you do not
|
|
80
|
+
want it to open a browser:
|
|
259
81
|
|
|
260
82
|
```sh
|
|
261
|
-
|
|
83
|
+
token-harness ui --no-open
|
|
262
84
|
```
|
|
263
85
|
|
|
264
|
-
|
|
265
|
-
HarnessTrim, or a coding agent.
|
|
266
|
-
|
|
267
|
-
### Managed compatibility rows
|
|
268
|
-
|
|
269
|
-
Token Harness changes a harness configuration only when a reviewed compatibility row covers the
|
|
270
|
-
exact provider version, harness version, platform, and configuration schema. Four rows ship, and each
|
|
271
|
-
names the recording it stands on:
|
|
272
|
-
|
|
273
|
-
| Provider | Harness | Platform | Tested versions | Tier |
|
|
274
|
-
| --- | --- | --- | --- | --- |
|
|
275
|
-
| RTK | Claude Code | Windows | rtk 0.44.0, Claude Code 2.1.220 | `canary` |
|
|
276
|
-
| HarnessTrim | Claude Code | Windows | harnesstrim 0.1.0, Claude Code 2.1.220 | `config-only` |
|
|
277
|
-
| HarnessTrim | Codex | Windows | harnesstrim 0.1.0, Codex 0.146.0 | `config-only` |
|
|
278
|
-
| HarnessTrim | Codex | Linux (non-WSL) | harnesstrim 0.2.1, Codex 0.152.1 | `config-only` |
|
|
279
|
-
|
|
280
|
-
Everything else is refused, and that is the design rather than a gap: `doctor` detects and reports on
|
|
281
|
-
every supported platform, and only the *mutation* is narrower. An uncovered combination exits 9 and
|
|
282
|
-
the diagnostic names what is missing — the reviewed fixture, or the nearest row it does have.
|
|
283
|
-
|
|
284
|
-
A live Linux recording on 2026-09-02 promoted exactly one combination to managed mutation:
|
|
285
|
-
Codex 0.152.1 + HarnessTrim 0.2.1 on non-WSL Linux. The fixture covers an empty state, brownfield
|
|
286
|
-
user-owned files, skills-only apply, drift, verified rollback, and surgical uninstall. Nearby Codex
|
|
287
|
-
or HarnessTrim versions and WSL remain refused.
|
|
288
|
-
|
|
289
|
-
What is not covered today, and why:
|
|
290
|
-
|
|
291
|
-
- **macOS, WSL, and other Linux version combinations.** The recordings a row needs are states of a
|
|
292
|
-
real machine. Only the exact Linux combination above has been reviewed, so all other combinations
|
|
293
|
-
continue to refuse managed mutation while detection, verification, and measurement still work.
|
|
294
|
-
- **OpenCode, and permanently rather than pending.** Both providers are detected, adopted, verified
|
|
295
|
-
and measured there, and neither is written. RTK reaches OpenCode through a plugin module its own
|
|
296
|
-
installer places globally, which this build has no action for. HarnessTrim's OpenCode installer
|
|
297
|
-
writes a plugin wrapper *and runs an npm install*, so a containment boundary covering what it wrote
|
|
298
|
-
would hold a `node_modules` tree — and that is not a decision deferred for want of a fixture. A
|
|
299
|
-
dependency tree is not configuration, so it cannot be a reviewed write set; snapshotting it on
|
|
300
|
-
every apply to keep the rollback honest would be slow and would be restoring upstream's install
|
|
301
|
-
rather than our change; and excluding it would leave a transaction claiming a reversibility it does
|
|
302
|
-
not have. So the assignment is not producible, and RFC 0003 is explicit about what that means: a
|
|
303
|
-
capability the provider has but cannot be asked for is not an assignable capability. OpenCode stays
|
|
304
|
-
adoption-only by decision.
|
|
305
|
-
- **RTK on Codex.** Not managed, and no row: RTK writes a Claude-shaped hook list and nothing else.
|
|
306
|
-
- **A newer Claude Code.** The range is a single observed version. `2.1.221` reads `unknown-newer` and
|
|
307
|
-
refuses rather than assuming it behaves like `2.1.220`.
|
|
308
|
-
|
|
309
|
-
The recordings are under `tests/fixtures/rows/`, one directory per row, each with a README stating
|
|
310
|
-
which stages exist and which do not.
|
|
311
|
-
|
|
312
|
-
## How the components fit together
|
|
313
|
-
|
|
314
|
-
There are three separate layers. Installing one does not automatically provide the others.
|
|
315
|
-
|
|
316
|
-
| Layer | Examples | Who installs it? |
|
|
317
|
-
| --- | --- | --- |
|
|
318
|
-
| Coding agent (harness) | Claude Code, Codex, OpenCode, Hermes, Pi | You, using the agent's official installer |
|
|
319
|
-
| Token Harness | `token-harness` | You, from npm or this repository |
|
|
320
|
-
| Optimization provider | RTK, HarnessTrim | Both can be installed by Token Harness where a compatibility row covers the combination; otherwise install them with their own installers and Token Harness detects and measures them |
|
|
321
|
-
|
|
322
|
-
Token Harness does not install Claude Code, Codex, OpenCode, Hermes, or Pi. Install and run at least one of
|
|
323
|
-
them first so that `token-harness doctor` can detect it.
|
|
324
|
-
|
|
325
|
-
| Provider | Claude Code | Codex | OpenCode | Hermes | Pi | Installed by Token Harness |
|
|
326
|
-
| --- | --- | --- | --- | --- | --- | --- |
|
|
327
|
-
| RTK | Configure, verify, and measure | Not managed | Detect, adopt, verify, and measure | Not managed | Not managed | **Yes**, for the supported Claude Code path |
|
|
328
|
-
| HarnessTrim | Claude skills only; no reducer hook or reduce-pipe instruction | Detect, adopt, verify, and measure | Detect, adopt, verify, and measure | Detect, verify, and measure | Detect, verify, and measure | **Yes**, on a covered row — see above |
|
|
329
|
-
|
|
330
|
-
"Not managed" does not mean the upstream tool cannot support that agent. It means this release
|
|
331
|
-
does not claim ownership of that integration and will not modify it.
|
|
332
|
-
|
|
333
|
-
Hermes is read-only in both directions: the adapter finds the HarnessTrim plugin, reads whether it
|
|
334
|
-
is enabled, and imports the telemetry it writes to `~/.hermes/harnesstrim-metrics.jsonl`, but nothing
|
|
335
|
-
here enables the plugin or restarts the gateway. Enabling it is
|
|
336
|
-
`hermes plugins enable harnesstrim`, and that stays your command to run. No compatibility row ships
|
|
337
|
-
for Hermes because a row is the precondition for a *mutation*, and none is proposed.
|
|
86
|
+
## Daily use
|
|
338
87
|
|
|
339
|
-
|
|
340
|
-
directories Pi auto-loads (`~/.pi/agent/extensions/` and `<project>/.pi/extensions/`) and verifies
|
|
341
|
-
the configuration, but nothing here installs it, and nothing here can say which mode it runs in —
|
|
342
|
-
the extension defaults to `dryrun` and only `HARNESSTRIM_MODE=active` in Pi's environment makes it
|
|
343
|
-
reduce. Installing it is `harnesstrim install pi --apply`, and that stays your command to run. No
|
|
344
|
-
compatibility row ships for Pi because a row is the precondition for a *mutation*, and none is
|
|
345
|
-
proposed.
|
|
88
|
+
There is no mandatory command loop. These are tools you use when they answer a question:
|
|
346
89
|
|
|
347
|
-
|
|
348
|
-
|
|
349
|
-
|
|
350
|
-
|
|
351
|
-
|
|
352
|
-
|
|
353
|
-
|
|
354
|
-
|
|
355
|
-
## Installing each component
|
|
356
|
-
|
|
357
|
-
### 1. Install Token Harness
|
|
90
|
+
| When you want to know... | Run |
|
|
91
|
+
| --- | --- |
|
|
92
|
+
| Is everything still connected? | `token-harness ui` |
|
|
93
|
+
| What should I do before a demanding task? | `token-harness optimize --task hard --profile quality` |
|
|
94
|
+
| Is an integration actually working? | `token-harness verify` |
|
|
95
|
+
| How much reducer saving has been measured? | `token-harness metrics --since 7d` |
|
|
96
|
+
| Are safer provider updates available? | `token-harness update` |
|
|
358
97
|
|
|
359
|
-
|
|
98
|
+
If a command finishes with **no action required**, stop there and use your coding agent
|
|
99
|
+
normally. Token Harness should not send you around a `ui → optimize → ui` loop.
|
|
360
100
|
|
|
361
|
-
|
|
362
|
-
npm install --global token-harness
|
|
363
|
-
token-harness --help
|
|
364
|
-
```
|
|
101
|
+
## Ask your AI to install Token Harness
|
|
365
102
|
|
|
366
|
-
|
|
103
|
+
You can give this prompt to Claude Code or Codex:
|
|
367
104
|
|
|
368
|
-
```
|
|
369
|
-
npm
|
|
105
|
+
```text
|
|
106
|
+
Install the latest stable Token Harness from npm on this computer, then run
|
|
107
|
+
`token-harness setup`. Do not install or replace Claude Code, Codex, or any
|
|
108
|
+
optimization provider unless Token Harness's supported plan explicitly requires it.
|
|
109
|
+
|
|
110
|
+
Explain the setup result in plain language: what was detected, what already works,
|
|
111
|
+
what would change, and the single next step. Do not expose credentials, cookies,
|
|
112
|
+
tokens, raw home paths, or private project contents. If setup proposes a supported
|
|
113
|
+
configuration change, show me the short plan and ask before running
|
|
114
|
+
`token-harness setup --yes`. After an approved change, verify it and open
|
|
115
|
+
`token-harness ui` once. Then tell me clearly that setup is complete and that I should
|
|
116
|
+
continue using my normal coding-agent command. Do not invent additional Token Harness
|
|
117
|
+
steps when no action is required.
|
|
370
118
|
```
|
|
371
119
|
|
|
372
|
-
|
|
373
|
-
|
|
374
|
-
#### Build and install from source
|
|
120
|
+
The AI should ask before the `--yes` step because that is the point where coding-agent
|
|
121
|
+
configuration may change.
|
|
375
122
|
|
|
376
|
-
|
|
123
|
+
## What normal output looks like
|
|
377
124
|
|
|
378
|
-
|
|
379
|
-
git clone https://github.com/giuliastro/token-harness.git
|
|
380
|
-
cd token-harness
|
|
381
|
-
corepack enable
|
|
382
|
-
pnpm install
|
|
383
|
-
pnpm build
|
|
384
|
-
pnpm package
|
|
385
|
-
npm install --global ./dist/package
|
|
386
|
-
token-harness --version
|
|
387
|
-
```
|
|
125
|
+
A healthy final check is intentionally short:
|
|
388
126
|
|
|
389
|
-
|
|
390
|
-
|
|
391
|
-
If `corepack` is unavailable, install the pinned package manager with
|
|
392
|
-
`npm install --global pnpm@10.33.4` instead.
|
|
127
|
+
```text
|
|
128
|
+
TOKEN HARNESS - READY
|
|
393
129
|
|
|
394
|
-
|
|
130
|
+
WHAT WORKS
|
|
131
|
+
Codex: configured (0.146.0)
|
|
132
|
+
HarnessTrim: active on Codex
|
|
395
133
|
|
|
396
|
-
|
|
134
|
+
CHANGES
|
|
135
|
+
Nothing changed.
|
|
397
136
|
|
|
398
|
-
|
|
399
|
-
|
|
137
|
+
NEXT STEP
|
|
138
|
+
Use your coding agent normally; configured optimizers run automatically.
|
|
400
139
|
```
|
|
401
140
|
|
|
402
|
-
|
|
141
|
+
A newer-than-tested combination is not presented as if the whole setup were broken:
|
|
403
142
|
|
|
404
|
-
|
|
405
|
-
|
|
406
|
-
|
|
407
|
-
The channel selected by this release is:
|
|
408
|
-
|
|
409
|
-
| Platform | Channel used by the plan | Required command on `PATH` |
|
|
410
|
-
| --- | --- | --- |
|
|
411
|
-
| Windows | WinGet package `rtk-ai.rtk` | `winget` |
|
|
412
|
-
| macOS | Cargo package `rtk` | `cargo` |
|
|
413
|
-
| Linux and WSL | Cargo package `rtk` | `cargo` |
|
|
414
|
-
|
|
415
|
-
The Cargo path in this release invokes `cargo install rtk`. That channel is declared but has not
|
|
416
|
-
been exercised by this project, and upstream documents a crates.io name collision. On macOS,
|
|
417
|
-
Linux, and WSL, the safer current route is to install RTK with an upstream-recommended method,
|
|
418
|
-
confirm that `rtk gain` works, and let Token Harness adopt and configure the existing binary.
|
|
143
|
+
```text
|
|
144
|
+
TOKEN HARNESS - READY WITH LIMITATIONS
|
|
419
145
|
|
|
420
|
-
|
|
146
|
+
WHAT WORKS
|
|
147
|
+
Claude Code: configured
|
|
148
|
+
RTK: active on Claude Code
|
|
421
149
|
|
|
422
|
-
|
|
423
|
-
token-harness
|
|
150
|
+
NEXT STEP
|
|
151
|
+
token-harness verify
|
|
152
|
+
You can keep working; verify the active integrations when convenient.
|
|
424
153
|
```
|
|
425
154
|
|
|
426
|
-
|
|
427
|
-
rewriting it. User-owned configuration remains user-owned.
|
|
428
|
-
|
|
429
|
-
Important boundaries:
|
|
430
|
-
|
|
431
|
-
- Token Harness writes the reviewed hook itself; it does not run `rtk init`.
|
|
432
|
-
- A package install is not reversed by file rollback. `rollback` restores configuration files,
|
|
433
|
-
not installed binaries.
|
|
434
|
-
- `uninstall` removes only integration entries written by Token Harness; it deliberately leaves
|
|
435
|
-
the RTK executable installed.
|
|
436
|
-
- On native Windows, Claude Code exposes both Bash and PowerShell tool families. The current RTK
|
|
437
|
-
matcher covers Bash only, so `doctor` can correctly report PowerShell as bypassed.
|
|
438
|
-
|
|
439
|
-
For manual installation or use outside Token Harness's managed surface, follow the
|
|
440
|
-
[RTK installation guide](https://github.com/rtk-ai/rtk/blob/master/INSTALL.md), then run:
|
|
155
|
+
Need the evidence behind a summary? Add `--verbose`:
|
|
441
156
|
|
|
442
157
|
```sh
|
|
443
|
-
|
|
444
|
-
rtk gain
|
|
445
|
-
token-harness doctor --provider rtk
|
|
158
|
+
token-harness doctor --verbose
|
|
446
159
|
```
|
|
447
160
|
|
|
448
|
-
|
|
449
|
-
`rtk`.
|
|
450
|
-
|
|
451
|
-
### 3. Install or adopt HarnessTrim
|
|
452
|
-
|
|
453
|
-
With HarnessTrim on `PATH`, `token-harness plan --harness claude` can install its Claude skills
|
|
454
|
-
without creating the competing Bash hook or reduce-pipe instruction. The invocation it delegates to
|
|
455
|
-
is:
|
|
161
|
+
Need stable machine-readable output for automation? Add `--json`:
|
|
456
162
|
|
|
457
163
|
```sh
|
|
458
|
-
|
|
164
|
+
token-harness doctor --json
|
|
165
|
+
token-harness ui --json
|
|
459
166
|
```
|
|
460
167
|
|
|
461
|
-
|
|
462
|
-
first as a dry run and then with its explicit apply flag. Consult the
|
|
463
|
-
[HarnessTrim README](https://github.com/giuliastro/HarnessTrim#quick-start) because its adapter
|
|
464
|
-
contents, modes, and telemetry differ by coding agent.
|
|
168
|
+
`--json` keeps the complete schema-1 result and diagnostics; it is not shortened.
|
|
465
169
|
|
|
466
|
-
|
|
170
|
+
## Three commands to remember
|
|
467
171
|
|
|
468
172
|
```sh
|
|
469
|
-
token-harness
|
|
470
|
-
token-harness
|
|
471
|
-
token-harness
|
|
472
|
-
token-harness metrics --provider harnesstrim --since 7d
|
|
173
|
+
token-harness setup
|
|
174
|
+
token-harness ui
|
|
175
|
+
token-harness optimize
|
|
473
176
|
```
|
|
474
177
|
|
|
475
|
-
|
|
476
|
-
|
|
477
|
-
|
|
478
|
-
|
|
479
|
-
|
|
480
|
-
verification can still inspect configuration, but `metrics` has no HarnessTrim events to import.
|
|
178
|
+
| Command | Answer |
|
|
179
|
+
| --- | --- |
|
|
180
|
+
| `setup` | Is Token Harness ready, and what is my one next step? |
|
|
181
|
+
| `ui` | What is active, how much allowance is visible, and do I need to do anything? |
|
|
182
|
+
| `optimize` | What is the best evidence-based action for the task I am starting? |
|
|
481
183
|
|
|
482
|
-
|
|
483
|
-
intercepts per coding agent, the flags that narrow an install, and the paths each install writes.
|
|
184
|
+
For example:
|
|
484
185
|
|
|
485
186
|
```sh
|
|
486
|
-
|
|
187
|
+
token-harness optimize --task hard --profile quality
|
|
188
|
+
token-harness optimize --task mechanical --profile economy
|
|
487
189
|
```
|
|
488
190
|
|
|
489
|
-
|
|
490
|
-
upstream change is reported rather than assumed compatible. A disagreement becomes a
|
|
491
|
-
`provider-capabilities-drift` warning naming both sides. A build older than the command cannot
|
|
492
|
-
answer; Token Harness then falls back to its own recorded declaration and reports nothing, because a
|
|
493
|
-
provider that cannot be asked must still be describable.
|
|
191
|
+
## Safety and privacy
|
|
494
192
|
|
|
495
|
-
|
|
193
|
+
Token Harness is conservative by design:
|
|
496
194
|
|
|
497
|
-
|
|
195
|
+
- normal read-only commands do not change coding-agent or project configuration;
|
|
196
|
+
- `setup --yes`, `apply --yes`, `update --yes`, `rollback --yes`, and
|
|
197
|
+
`uninstall --yes` are the explicit configuration-changing forms;
|
|
198
|
+
- plans are checked again immediately before they are applied;
|
|
199
|
+
- existing files are backed up before a managed write;
|
|
200
|
+
- only exact Token Harness-owned entries are removed by `uninstall`;
|
|
201
|
+
- newer or untested combinations are reported, not guessed;
|
|
202
|
+
- an available provider update outside reviewed compatibility is kept out rather than
|
|
203
|
+
forced, and the installed working version stays in place;
|
|
204
|
+
- the dashboard binds only to the local loopback address and provides no mutation API;
|
|
205
|
+
- source code, prompts, command contents, credentials, and cookies are not sent to a
|
|
206
|
+
Token Harness service.
|
|
498
207
|
|
|
499
|
-
|
|
500
|
-
|
|
501
|
-
|
|
502
|
-
|
|
503
|
-
This answers:
|
|
208
|
+
Plans, receipts, metrics, and backups stay in the local Token Harness state directory.
|
|
209
|
+
See [RFC 0004](docs/rfcs/0004-safety-and-installation.md) for the execution model and
|
|
210
|
+
[RFC 0006](docs/rfcs/0006-cli-contract.md) for CLI/JSON guarantees.
|
|
504
211
|
|
|
505
|
-
|
|
506
|
-
- which providers are installed and runnable;
|
|
507
|
-
- which agent configuration files exist;
|
|
508
|
-
- which provider is wired to which agent;
|
|
509
|
-
- whether Token Harness owns the integration or merely adopted it;
|
|
510
|
-
- whether a version, configuration file, or tool-family matcher needs attention;
|
|
511
|
-
- whether the installed provider's own capability declaration still agrees with the one Token
|
|
512
|
-
Harness records.
|
|
212
|
+
## Supported optimizations
|
|
513
213
|
|
|
514
|
-
|
|
515
|
-
|
|
516
|
-
| State | Meaning |
|
|
517
|
-
| --- | --- |
|
|
518
|
-
| `not found` / `absent` | The executable and usable configuration were not detected |
|
|
519
|
-
| `installed` | The provider runs but is not connected to a supported agent |
|
|
520
|
-
| `configured` | A relevant hook or plugin entry exists |
|
|
521
|
-
| `broken` | Configuration refers to something missing or unreadable |
|
|
522
|
-
| `set up by you` | Token Harness adopted existing configuration and will not remove it |
|
|
523
|
-
| `set up by this tool` | A committed Token Harness transaction owns the exact entry |
|
|
214
|
+
Token Harness can detect and measure several independent local tools:
|
|
524
215
|
|
|
525
|
-
|
|
216
|
+
| Provider | Purpose | Management |
|
|
217
|
+
| --- | --- | --- |
|
|
218
|
+
| [RTK](https://github.com/rtk-ai/rtk) | Shell-command rewriting and output reduction | Managed only for reviewed combinations |
|
|
219
|
+
| [HarnessTrim](https://github.com/giuliastro/HarnessTrim) | Deterministic reducers and harness adapters | Managed only for reviewed combinations |
|
|
220
|
+
| [cclimits](https://github.com/cruzanstx/cclimits) | Optional live/local quota companion | Read-only; never installed automatically |
|
|
221
|
+
| [ccusage](https://github.com/ccusage/ccusage) | Local usage history | Read-only; never installed automatically |
|
|
526
222
|
|
|
527
|
-
|
|
223
|
+
A provider you installed yourself remains yours. Token Harness can adopt observable
|
|
224
|
+
configuration without claiming ownership of the executable.
|
|
528
225
|
|
|
529
|
-
|
|
530
|
-
|
|
531
|
-
|
|
226
|
+
Exact reviewed provider/harness/platform/version combinations are generated in
|
|
227
|
+
[docs/matrices.md](docs/matrices.md). A combination outside that table can still be
|
|
228
|
+
detected and inspected, but Token Harness will not mutate it.
|
|
532
229
|
|
|
533
|
-
|
|
230
|
+
## Advanced commands
|
|
534
231
|
|
|
535
|
-
|
|
536
|
-
token-harness plan --harness claude
|
|
537
|
-
token-harness plan --provider rtk
|
|
538
|
-
token-harness plan --project /path/to/project
|
|
539
|
-
```
|
|
232
|
+
Most people do not need this section. Run `token-harness <command> --help` for details.
|
|
540
233
|
|
|
541
|
-
|
|
234
|
+
| Command | Purpose | Changes agent/project config? |
|
|
235
|
+
| --- | --- | --- |
|
|
236
|
+
| `doctor` | Detect harnesses, providers, versions, and problems | No |
|
|
237
|
+
| `budget` | Read authoritative/reported allowance windows | No |
|
|
238
|
+
| `context` | Inspect model settings, instructions, and MCP exposure | No |
|
|
239
|
+
| `mcp` | Focus on MCP server/tool health | No |
|
|
240
|
+
| `history` | Summarize local usage through an installed ccusage | No |
|
|
241
|
+
| `plan` | Prepare exact supported changes | No; stores local plan state |
|
|
242
|
+
| `apply` | Apply a reviewed stored plan | Yes, only with `--yes` |
|
|
243
|
+
| `verify` | Check the declared integration tier | No |
|
|
244
|
+
| `metrics` | Report attributable reducer savings | No |
|
|
245
|
+
| `status` | Report pipelines, drift, and importer modes | No |
|
|
246
|
+
| `update` | Check/update installed providers; unreviewed targets stay installed | Yes, only with `--yes` |
|
|
247
|
+
| `rollback` | Restore the latest transaction snapshot | Yes, only with `--yes` |
|
|
248
|
+
| `uninstall` | Remove owned integration entries | Yes, only with `--yes` |
|
|
249
|
+
| `schedule` | Compare Claude Code and Codex using available evidence | No |
|
|
250
|
+
| `handoff` | Build a bounded cross-harness handoff | No |
|
|
251
|
+
| `benchmark*`, `transfer*` | Capture and compare empirical evidence | Local state only |
|
|
542
252
|
|
|
543
|
-
|
|
544
|
-
- `Excluded`: detected providers intentionally left out;
|
|
545
|
-
- `Actions`: every package operation and file change;
|
|
546
|
-
- `Network`: destinations contacted by later mutation;
|
|
547
|
-
- `Elevation`: whether administrator/root access would be required;
|
|
548
|
-
- `Backups`: how many files will be snapshotted.
|
|
253
|
+
## Troubleshooting
|
|
549
254
|
|
|
550
|
-
`
|
|
551
|
-
Token Harness's private state directory so the exact reviewed artifact can be applied later.
|
|
255
|
+
### `token-harness` is not found
|
|
552
256
|
|
|
553
|
-
|
|
257
|
+
Check that Node is new enough and the package is installed:
|
|
554
258
|
|
|
555
259
|
```sh
|
|
556
|
-
|
|
260
|
+
node --version
|
|
261
|
+
npm list --global token-harness
|
|
557
262
|
```
|
|
558
263
|
|
|
559
|
-
|
|
560
|
-
|
|
264
|
+
Node must be at least 22.13. Reopen the terminal after installation if needed.
|
|
265
|
+
|
|
266
|
+
### Setup needs attention
|
|
561
267
|
|
|
562
|
-
|
|
268
|
+
Run the single command it prints. For technical evidence:
|
|
563
269
|
|
|
564
270
|
```sh
|
|
565
|
-
token-harness
|
|
271
|
+
token-harness doctor --verbose
|
|
566
272
|
```
|
|
567
273
|
|
|
568
|
-
|
|
569
|
-
|
|
570
|
-
triggers automatic restoration and the result states whether that restoration was verified.
|
|
274
|
+
Do not force an unsupported plan. Open an issue with the redacted `--json` result if
|
|
275
|
+
you believe the combination should be supported.
|
|
571
276
|
|
|
572
|
-
|
|
277
|
+
### `update` finds a newer version but keeps the installed one
|
|
573
278
|
|
|
574
|
-
|
|
279
|
+
That is normally a safety decision, not a failed installation. Token Harness found a
|
|
280
|
+
newer provider release but does not yet have reviewed compatibility evidence for the
|
|
281
|
+
active provider × harness × platform combination. Keep using the installed version; no
|
|
282
|
+
manual upgrade is required.
|
|
575
283
|
|
|
576
|
-
|
|
577
|
-
Open the configured coding agent and ask it to run a normal shell command such as `git status` or
|
|
578
|
-
a test command. Then return to the terminal.
|
|
284
|
+
### Verification says `not-exercised`
|
|
579
285
|
|
|
580
|
-
|
|
581
|
-
|
|
582
|
-
Use both commands; they answer different questions:
|
|
286
|
+
Restart the coding agent, use it for one normal command, and run:
|
|
583
287
|
|
|
584
288
|
```sh
|
|
585
|
-
token-harness status
|
|
586
289
|
token-harness verify
|
|
587
290
|
```
|
|
588
291
|
|
|
589
|
-
|
|
590
|
-
|
|
591
|
-
|
|
592
|
-
`verify` checks the strongest evidence the integration declares:
|
|
292
|
+
No observed operation is different from a failed integration, so Token Harness reports
|
|
293
|
+
the two states separately.
|
|
593
294
|
|
|
594
|
-
|
|
595
|
-
| --- | --- |
|
|
596
|
-
| `presence` | The executable resolves and reports a version |
|
|
597
|
-
| `config-only` | The expected configuration entry exists |
|
|
598
|
-
| `canary` | Provider records show a real operation crossed the interception point |
|
|
599
|
-
|
|
600
|
-
`config-only` is not proof that the hook ran. It is the honest ceiling for integrations whose
|
|
601
|
-
runtime state cannot be observed externally.
|
|
602
|
-
|
|
603
|
-
`not-exercised` means no attributable operation has been observed yet. It is neither success nor
|
|
604
|
-
failure: run a command through the agent and check again.
|
|
295
|
+
## Updating or undoing
|
|
605
296
|
|
|
606
|
-
|
|
297
|
+
Update the CLI:
|
|
607
298
|
|
|
608
299
|
```sh
|
|
609
|
-
token-harness
|
|
610
|
-
token-harness
|
|
611
|
-
token-harness metrics --since 2026-07-01 --until 2026-08-01
|
|
612
|
-
token-harness metrics --provider rtk --since 7d
|
|
300
|
+
npm install --global token-harness@latest
|
|
301
|
+
token-harness setup
|
|
613
302
|
```
|
|
614
303
|
|
|
615
|
-
|
|
616
|
-
accepted. Date boundaries are midnight UTC.
|
|
617
|
-
|
|
618
|
-
The report keeps measurement types and units separate:
|
|
619
|
-
|
|
620
|
-
| Report line | Interpretation |
|
|
621
|
-
| --- | --- |
|
|
622
|
-
| `Exact local` | Before and after token counts were observed for the same operation |
|
|
623
|
-
| `Estimated local` | The payload changed, but the reported unit or tokenizer is an estimate |
|
|
624
|
-
| `Counterfactual` | A dry run measured what could have changed; it is not realized saving |
|
|
625
|
-
| `End-to-end billed` | Comparable billed sessions were measured; otherwise it says `no A/B run` |
|
|
626
|
-
| `Coverage` | Share of relevant operations that were actually changed |
|
|
627
|
-
| `Bypassed` | Operations observed but passed through unchanged or outside coverage |
|
|
628
|
-
|
|
629
|
-
Token counts are never added to character counts, and estimated or counterfactual values are never
|
|
630
|
-
silently merged into an exact total.
|
|
631
|
-
|
|
632
|
-
The report covers one project: the one `--project` names, or the current directory. An operation a
|
|
633
|
-
provider recorded without a directory belongs to no project and is excluded, with a count reported
|
|
634
|
-
so the difference is reconcilable. When no project identity can be established the report says so
|
|
635
|
-
rather than presenting every project's events as one project's figures.
|
|
636
|
-
|
|
637
|
-
## Undoing changes
|
|
638
|
-
|
|
639
|
-
Choose the command based on what you want to undo:
|
|
304
|
+
Remove only Token Harness-owned integration entries:
|
|
640
305
|
|
|
641
306
|
```sh
|
|
642
|
-
# Remove only exact integration entries owned by Token Harness.
|
|
643
307
|
token-harness uninstall --yes
|
|
644
|
-
|
|
645
|
-
# Restore all files from the most recent committed transaction snapshot.
|
|
646
|
-
token-harness rollback --yes
|
|
647
|
-
```
|
|
648
|
-
|
|
649
|
-
`uninstall` is usually the safer choice after subsequent manual edits: it is surgical and refuses
|
|
650
|
-
to remove an owned entry if its content no longer matches what Token Harness wrote.
|
|
651
|
-
|
|
652
|
-
`rollback` restores whole files to their pre-transaction bytes. Changes made to those files after
|
|
653
|
-
the transaction are therefore also reverted. It does not restore or remove provider packages.
|
|
654
|
-
|
|
655
|
-
Neither command removes user-owned RTK or HarnessTrim configuration.
|
|
656
|
-
|
|
657
|
-
## Command reference
|
|
658
|
-
|
|
659
|
-
| Command | Purpose | Changes agent/project configuration? |
|
|
660
|
-
| --- | --- | --- |
|
|
661
|
-
| `doctor` | Detect agents, providers, ownership, versions, and problems | No |
|
|
662
|
-
| `budget` | Read live quota/headroom windows where the harness exposes them | No |
|
|
663
|
-
| `context` | Inspect effective model/config, instructions, MCP exposure, and tool inventory | No |
|
|
664
|
-
| `mcp` | Inspect MCP servers and tool-schema exposure | No |
|
|
665
|
-
| `history` | Summarize attributable local usage history | No |
|
|
666
|
-
| `optimize` | Combine quota pacing, context pressure, and task/profile policy into advice | No |
|
|
667
|
-
| `schedule` | Recommend Claude Code or Codex from independently attributable evidence | No |
|
|
668
|
-
| `handoff` | Build a bounded compact handoff for an in-progress cross-harness move | No |
|
|
669
|
-
| `plan` | Resolve ownership and preview exact actions; use `--native-policy` for supported harness-native changes | No; stores the plan in private state |
|
|
670
|
-
| `apply` | Apply a reviewed plan transactionally | Yes, only with `--yes` |
|
|
671
|
-
| `status` | Detect drift and competing hooks | No |
|
|
672
|
-
| `verify` | Check the declared verification tier | No |
|
|
673
|
-
| `metrics` | Import provider records and report savings | No; updates only Token Harness state |
|
|
674
|
-
| `benchmark` | Compare an explicit baseline/optimized receipt pair | No |
|
|
675
|
-
| `benchmark-start` | Start an empirical task capture | No agent/project config change; records Token Harness benchmark state |
|
|
676
|
-
| `benchmark-finish` | Finish an empirical task capture and record the outcome | No agent/project config change; records Token Harness benchmark state |
|
|
677
|
-
| `benchmark-matrix` | Aggregate complete project-scoped benchmark pairs by task class and evidence | No |
|
|
678
|
-
| `transfer` | Evaluate one empirical cross-harness benchmark pair and exact handoff | No |
|
|
679
|
-
| `transfer-record` | Persist one immutable project-scoped transfer evidence receipt | No agent/project config change; records Token Harness benchmark state |
|
|
680
|
-
| `update` | Query channels and update installed providers | Yes, only with `--yes` |
|
|
681
|
-
| `rollback` | Restore files from the latest committed transaction | Yes, only with `--yes` |
|
|
682
|
-
| `uninstall` | Remove owned integration entries | Yes, only with `--yes` |
|
|
683
|
-
|
|
684
|
-
Every command supports `--help`. Common filters are:
|
|
685
|
-
|
|
686
|
-
```text
|
|
687
|
-
--harness <id>
|
|
688
|
-
--provider <id>
|
|
689
|
-
--project <directory>
|
|
690
|
-
--task mechanical|standard|hard|critical
|
|
691
|
-
--profile economy|balanced|quality|custom
|
|
692
|
-
--reserve <percent>
|
|
693
|
-
--native-policy
|
|
694
|
-
--plan <plan-id>
|
|
695
|
-
--json
|
|
696
|
-
```
|
|
697
|
-
|
|
698
|
-
## Automation and JSON output
|
|
699
|
-
|
|
700
|
-
Use `--json` in scripts:
|
|
701
|
-
|
|
702
|
-
```sh
|
|
703
|
-
token-harness doctor --json
|
|
704
|
-
token-harness verify --json
|
|
705
|
-
token-harness metrics --since 7d --json
|
|
706
308
|
```
|
|
707
309
|
|
|
708
|
-
|
|
709
|
-
|
|
710
|
-
```json
|
|
711
|
-
{
|
|
712
|
-
"schemaVersion": 1,
|
|
713
|
-
"command": "verify",
|
|
714
|
-
"toolVersion": "0.1.0",
|
|
715
|
-
"status": "ok",
|
|
716
|
-
"exitCode": 0,
|
|
717
|
-
"data": {},
|
|
718
|
-
"diagnostics": []
|
|
719
|
-
}
|
|
720
|
-
```
|
|
721
|
-
|
|
722
|
-
Important exit codes:
|
|
723
|
-
|
|
724
|
-
| Code | Meaning |
|
|
725
|
-
| ---: | --- |
|
|
726
|
-
| 0 | Completed with nothing actionable |
|
|
727
|
-
| 2 | Invalid command or argument |
|
|
728
|
-
| 3 | A read-only check found an actionable problem |
|
|
729
|
-
| 4 | A capability conflict blocks the plan |
|
|
730
|
-
| 5 | The environment drifted from the stored plan or journal |
|
|
731
|
-
| 6 | Mutation failed and rollback was verified |
|
|
732
|
-
| 7 | Mutation failed and state was not fully restored; inspect the named paths |
|
|
733
|
-
| 8 | The command needs explicit confirmation (`--yes`) |
|
|
734
|
-
| 9 | Unsupported or unverifiable environment |
|
|
735
|
-
|
|
736
|
-
Do not treat every non-zero code as the same failure. In particular, code 8 is the expected result
|
|
737
|
-
of previewing a mutating command without approval.
|
|
738
|
-
|
|
739
|
-
## State, backups, and privacy
|
|
740
|
-
|
|
741
|
-
Token Harness stores plans, journals, backups, receipts, import cursors, and normalized metrics
|
|
742
|
-
outside the repository:
|
|
743
|
-
|
|
744
|
-
| Platform | Default state root |
|
|
745
|
-
| --- | --- |
|
|
746
|
-
| Windows | `%LOCALAPPDATA%\TokenHarness` |
|
|
747
|
-
| macOS | `~/Library/Application Support/TokenHarness` |
|
|
748
|
-
| Linux and WSL | `${XDG_STATE_HOME:-~/.local/state}/token-harness` |
|
|
749
|
-
|
|
750
|
-
Normalized metrics do not contain raw command text, tool output, source code, prompts, credentials,
|
|
751
|
-
or raw file paths. Provider records are read in place; Token Harness imports only normalized event
|
|
752
|
-
data.
|
|
753
|
-
|
|
754
|
-
## Troubleshooting
|
|
755
|
-
|
|
756
|
-
### `token-harness` is not found
|
|
757
|
-
|
|
758
|
-
Confirm Node and the global npm installation:
|
|
310
|
+
Restore complete files from the latest committed transaction snapshot:
|
|
759
311
|
|
|
760
312
|
```sh
|
|
761
|
-
|
|
762
|
-
npm list --global token-harness
|
|
763
|
-
npm prefix --global
|
|
313
|
+
token-harness rollback --yes
|
|
764
314
|
```
|
|
765
315
|
|
|
766
|
-
|
|
767
|
-
|
|
768
|
-
|
|
769
|
-
### `plan` says there is nothing to do
|
|
770
|
-
|
|
771
|
-
Run `token-harness doctor`. The usual causes are:
|
|
772
|
-
|
|
773
|
-
- no supported coding agent was detected;
|
|
774
|
-
- the requested provider does not claim that coding agent in this release;
|
|
775
|
-
- an existing user-managed integration already satisfies the target state;
|
|
776
|
-
- the safe profile excluded an overlapping provider.
|
|
316
|
+
`rollback` is whole-file time travel, so it can also revert later manual edits to those
|
|
317
|
+
files. Prefer `uninstall` when you only want to remove Token Harness-owned entries.
|
|
777
318
|
|
|
778
|
-
|
|
779
|
-
a `hooks` entry, which is Claude Code's schema — OpenCode's integration is a plugin module, so an
|
|
780
|
-
OpenCode scope produces no action and the existing installation is adopted instead. A Codex-only
|
|
781
|
-
machine produces no RTK action at all.
|
|
782
|
-
|
|
783
|
-
### The plan is blocked by `exclusive-scope-contested`
|
|
784
|
-
|
|
785
|
-
RTK and HarnessTrim both claim the same reducing surface. Token Harness will not choose an order or
|
|
786
|
-
overwrite either configuration. Remove or disable one integration using the tool that owns it, then
|
|
787
|
-
run `doctor` and `plan` again.
|
|
788
|
-
|
|
789
|
-
### `verify` reports `not-exercised`
|
|
790
|
-
|
|
791
|
-
Restart the coding agent, ask it to run a shell command through the configured tool family, then
|
|
792
|
-
run `token-harness verify` again. For a `config-only` integration, no stronger external receipt may
|
|
793
|
-
exist; the output states that limitation explicitly.
|
|
794
|
-
|
|
795
|
-
### `metrics` shows no data
|
|
796
|
-
|
|
797
|
-
Check all of the following:
|
|
798
|
-
|
|
799
|
-
- the provider has processed at least one operation in the requested time window;
|
|
800
|
-
- `rtk gain` works for RTK;
|
|
801
|
-
- HarnessTrim telemetry is enabled and `.harnesstrim/metrics.jsonl` exists for the project;
|
|
802
|
-
- `--project` points to the project whose records you expect;
|
|
803
|
-
- `--since` is not excluding older events.
|
|
804
|
-
|
|
805
|
-
An empty metrics report exits 0 because it is a valid observation, not a command failure.
|
|
806
|
-
|
|
807
|
-
### `doctor` or `status` reports `provider-capabilities-drift`
|
|
808
|
-
|
|
809
|
-
The installed provider's own capability declaration no longer agrees with the one Token Harness
|
|
810
|
-
records. The warning names both sides: what the recorded declaration claims, and what the installed
|
|
811
|
-
build reported. Nothing is modified, and the recorded declaration still drives planning.
|
|
812
|
-
|
|
813
|
-
Three disagreements are reported:
|
|
814
|
-
|
|
815
|
-
- a coding agent that Token Harness records a capability on is missing from the build's declaration;
|
|
816
|
-
- the reduction surface Token Harness records is absent from the surfaces the build reports;
|
|
817
|
-
- the build no longer covers a reviewed write-set path, or declares a path outside the reviewed
|
|
818
|
-
containment boundary.
|
|
819
|
-
|
|
820
|
-
The last one matters most before a delegated install. Rollback restores the reviewed boundary, so a
|
|
821
|
-
path outside it would survive a rollback. Re-review the write set at the installed version, or hold
|
|
822
|
-
at the reviewed one.
|
|
823
|
-
|
|
824
|
-
### A newer provider or agent version is reported
|
|
825
|
-
|
|
826
|
-
The tested ranges record versions actually exercised by this project. A newer version is reported
|
|
827
|
-
and handled conservatively rather than assumed compatible. Check [docs/matrices.md](docs/matrices.md)
|
|
828
|
-
and the upstream release notes before applying configuration changes.
|
|
829
|
-
|
|
830
|
-
## Development
|
|
319
|
+
## Develop from source
|
|
831
320
|
|
|
832
321
|
```sh
|
|
322
|
+
git clone https://github.com/giuliastro/token-harness.git
|
|
323
|
+
cd token-harness
|
|
833
324
|
corepack enable
|
|
834
325
|
pnpm install
|
|
835
326
|
pnpm typecheck
|
|
@@ -841,15 +332,10 @@ pnpm package
|
|
|
841
332
|
pnpm smoke:install
|
|
842
333
|
```
|
|
843
334
|
|
|
844
|
-
|
|
845
|
-
|
|
846
|
-
packed npm artifact.
|
|
847
|
-
|
|
848
|
-
Before changing architecture or public behavior, read [PLAN.md](PLAN.md) and the accepted RFCs in
|
|
849
|
-
[docs/rfcs](docs/rfcs). The CLI and JSON contract is defined by
|
|
850
|
-
[RFC 0006](docs/rfcs/0006-cli-contract.md).
|
|
335
|
+
Read [PLAN.md](PLAN.md) and the accepted [RFCs](docs/rfcs) before changing public
|
|
336
|
+
behavior or architecture.
|
|
851
337
|
|
|
852
338
|
## License
|
|
853
339
|
|
|
854
|
-
|
|
855
|
-
|
|
340
|
+
[Apache License 2.0](LICENSE). Referenced provider tools are independent projects with
|
|
341
|
+
their own licenses.
|