token-harness 0.1.25 → 0.1.27
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +132 -443
- package/package.json +1 -1
- package/sbom.json +3 -3
- package/token-harness.mjs +1027 -450
package/README.md
CHANGED
|
@@ -1,484 +1,194 @@
|
|
|
1
1
|
# Token Harness
|
|
2
2
|
|
|
3
|
-
**
|
|
3
|
+
**Make your Claude Code and Codex allowance go further. Measure what actually helps.**
|
|
4
4
|
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
5
|
+
[](https://www.npmjs.com/package/token-harness)
|
|
6
|
+
[](https://github.com/giuliastro/token-harness/actions/workflows/ci.yml)
|
|
7
|
+
[](LICENSE)
|
|
8
8
|
|
|
9
|
-
|
|
10
|
-
|
|
9
|
+
Coding agents spend context on long command output, repeated information and tool definitions.
|
|
10
|
+
Token Harness helps you reduce that overhead, connect compatible optimizers and see their measured
|
|
11
|
+
impact in one local dashboard. Keep working in Claude Code or Codex as usual.
|
|
11
12
|
|
|
12
|
-
|
|
13
|
+
For subscription users, the goal is **more accepted coding work within your included allowance**.
|
|
14
|
+
For API users, it is **less avoidable token usage with costs reported only when billing evidence is
|
|
15
|
+
available**. Token reduction, subscription quota and billed cost are different measurements; Token
|
|
16
|
+
Harness keeps them separate.
|
|
13
17
|
|
|
14
|
-
|
|
18
|
+
## Install and start
|
|
15
19
|
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
- the coding agent you want to use already signed in.
|
|
19
|
-
|
|
20
|
-
Install and open Token Harness:
|
|
20
|
+
You need **Node.js 22.13+** and an installed, signed-in **Claude Code or Codex**.
|
|
21
|
+
Run these commands in your terminal, including PowerShell on Windows:
|
|
21
22
|
|
|
22
23
|
```sh
|
|
23
24
|
npm install --global token-harness@latest
|
|
24
25
|
token-harness
|
|
25
26
|
```
|
|
26
27
|
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
The first screen is **Overview**. There is no separate Setup page to learn.
|
|
33
|
-
|
|
34
|
-
1. Token Harness detects Claude Code and Codex.
|
|
35
|
-
2. Each detected coding agent says either **Ready** or **Setup incomplete**.
|
|
36
|
-
3. If setup is incomplete, use the **Optimizer setup** matrix. It shows every optimizer against
|
|
37
|
-
every detected agent and lets you select one harness or both in the same review. The recommended
|
|
38
|
-
**RTK + HarnessTrim** baseline has its own action in that matrix and shows the exact safe plan
|
|
39
|
-
before anything changes.
|
|
40
|
-
4. Optional optimizers (mcptoon, GitNexus and Headroom) use the same matrix and per-optimizer action;
|
|
41
|
-
there is no repeated setup button under each coding-agent card.
|
|
42
|
-
5. **Health and updates** is maintenance, not another onboarding checklist. Normal setup performs its
|
|
43
|
-
own safety checks. Use **Re-check health** for troubleshooting and **Check for updates** to inspect
|
|
44
|
-
Token Harness and optimizer versions. If an update is available, the same dialog offers
|
|
45
|
-
**Install updates** after showing the versions.
|
|
46
|
-
6. Keep using Claude Code or Codex normally. Open **Results** when you want detailed evidence.
|
|
47
|
-
|
|
48
|
-
Opening the app does not change your configuration. A configuration or software change is always
|
|
49
|
-
previewed first and requires an explicit review and approval.
|
|
50
|
-
|
|
51
|
-
## The two views
|
|
52
|
-
|
|
53
|
-
### Overview
|
|
54
|
-
|
|
55
|
-
Overview is both the first-run screen and the normal status screen. It keeps the two product entities
|
|
56
|
-
separate:
|
|
57
|
-
|
|
58
|
-
- **Coding agents** — currently Claude Code and Codex. Detection only means Token Harness can see the
|
|
59
|
-
agent; first-run setup is complete for that agent only when both RTK and HarnessTrim are connected.
|
|
60
|
-
- **Optimizers** — RTK and HarnessTrim are the recommended baseline. mcptoon, GitNexus and Headroom
|
|
61
|
-
are optional reviewed integrations with narrower prerequisites and compatibility boundaries.
|
|
62
|
-
|
|
63
|
-
The status at the top always answers what to do next. States such as **Setup incomplete**,
|
|
64
|
-
**Installed · not connected**, **Needs attention**, or **Update available** have their action beside
|
|
65
|
-
the affected agent or optimizer instead of in a separate action list.
|
|
66
|
-
|
|
67
|
-
Overview also contains a compact **Measured impact** summary. Missing evidence is shown as unknown,
|
|
68
|
-
never as zero savings.
|
|
69
|
-
|
|
70
|
-
Advanced agent details and reasoning preferences are collapsed because they are not required for
|
|
71
|
-
first-run optimizer setup.
|
|
72
|
-
|
|
73
|
-
### Optimizer lifecycle and evidence
|
|
74
|
-
|
|
75
|
-
The recommended baseline remains **RTK + HarnessTrim** on individually reviewed combinations.
|
|
76
|
-
Token Harness also exposes **mcptoon**, **GitNexus** and **Headroom** as optional managed integrations
|
|
77
|
-
on exact reviewed lifecycle rows. Enabling an optional optimizer does not make it part of the
|
|
78
|
-
production baseline and does not create a savings claim.
|
|
79
|
-
|
|
80
|
-
Token Harness can prepare supported integration changes transactionally, show the exact plan, apply
|
|
81
|
-
it only after approval, verify what the declared tier can verify, and remove only configuration it
|
|
82
|
-
owns. GitNexus is never auto-indexed and its noncommercial license boundary stays visible. Headroom
|
|
83
|
-
remains config-only managed on its reviewed row; Token Harness does not bootstrap uv/Python or start
|
|
84
|
-
wrapper/proxy/deploy flows.
|
|
85
|
-
|
|
86
|
-
For maintainers, `token-harness stack-review --json` captures the exact configured provider versions
|
|
87
|
-
and managed harness sets for combined-stack review. Runtime evidence is credited only when it can be
|
|
88
|
-
attributed to the relevant harness; missing attribution remains unavailable rather than being copied
|
|
89
|
-
across rows.
|
|
90
|
-
|
|
91
|
-
Evaluation campaigns are an advanced maintainer workflow, not a novice setup step. Normal managed
|
|
92
|
-
setup stays in the unified **Optimization Stack** and historical candidate evidence never turns into
|
|
93
|
-
a savings, compatibility or promotion claim by itself.
|
|
94
|
-
|
|
95
|
-
### Results
|
|
96
|
-
|
|
97
|
-
Results keeps evidence separate from estimates. Depending on what is actually observable, it can
|
|
98
|
-
show:
|
|
99
|
-
|
|
100
|
-
- recorded optimizer output reduction;
|
|
101
|
-
- authoritative paired 5-hour / 7-day allowance evidence;
|
|
102
|
-
- API cost only when billed-token evidence and a verified price basis exist;
|
|
103
|
-
- paired quality evidence;
|
|
104
|
-
- experimental candidate evidence and campaign selection assessment when available;
|
|
105
|
-
- recent checks and changes from the current local app session.
|
|
106
|
-
|
|
107
|
-
**Not measured** means exactly that. Token Harness does not turn local token estimates into fake
|
|
108
|
-
subscription minutes, money or quota savings.
|
|
109
|
-
|
|
110
|
-
## Daily use
|
|
111
|
-
|
|
112
|
-
Keep launching `claude` or `codex` as usual. Deterministic optimizers that are installed, verified
|
|
113
|
-
and still beneficial are intended to remain enabled.
|
|
114
|
-
|
|
115
|
-
Token Harness does **not** need to stay open, does not need a permanent background daemon, and does
|
|
116
|
-
not need to decide before every command whether an optimizer should run.
|
|
28
|
+
A local browser dashboard opens. The browser app is the primary interface.
|
|
29
|
+
No Token Harness account or API key is required.
|
|
30
|
+
Your coding agent keeps its own authentication. Windows, macOS, Linux and WSL have explicit
|
|
31
|
+
compatibility checks; individual integrations may support a narrower set of versions/platforms.
|
|
117
32
|
|
|
118
|
-
|
|
119
|
-
verify integrations or re-evaluate the stack after a meaningful version/configuration change.
|
|
120
|
-
|
|
121
|
-
The app does not periodically reload the whole setup. A full read happens on initial open or when you
|
|
122
|
-
choose **Refresh**. Existing readings remain visible while a refresh runs. Applying a reviewed change
|
|
123
|
-
marks the displayed data as previous state instead of immediately launching another expensive full
|
|
124
|
-
read.
|
|
125
|
-
|
|
126
|
-
## What counts as savings
|
|
127
|
-
|
|
128
|
-
Token Harness keeps different evidence classes separate.
|
|
129
|
-
|
|
130
|
-
**Recorded output savings** are attributable reducer measurements. Providers, units and measurement
|
|
131
|
-
classes are not silently added together. Negative results and errors remain visible.
|
|
132
|
-
|
|
133
|
-
**5h / 7d allowance savings** require authoritative paired before/after allowance evidence. A
|
|
134
|
-
five-hour percentage may also be expressed as the equivalent share of that 300-minute allowance
|
|
135
|
-
window. Weekly quota is not converted into seven days of wall-clock compute.
|
|
136
|
-
|
|
137
|
-
**API cost** stays **Not measured yet** until attributable billed input/output tokens and a verified
|
|
138
|
-
model-price basis are available.
|
|
139
|
-
|
|
140
|
-
**Quality** is measured independently. A measured regression blocks a positive allowance-saving
|
|
141
|
-
claim rather than letting a smaller token number win by itself.
|
|
142
|
-
|
|
143
|
-
Upstream benchmark numbers are useful for deciding what to test; they are never copied directly into
|
|
144
|
-
your savings total.
|
|
145
|
-
|
|
146
|
-
For a terminal-only savings summary:
|
|
33
|
+
Prefer to try it before installing globally?
|
|
147
34
|
|
|
148
35
|
```sh
|
|
149
|
-
token-harness
|
|
36
|
+
npx --yes token-harness@latest
|
|
150
37
|
```
|
|
151
38
|
|
|
152
|
-
|
|
39
|
+
Updates for an `npx` launch use a new `npx` launch; automatic application updates require a verified
|
|
40
|
+
global npm installation.
|
|
153
41
|
|
|
154
|
-
##
|
|
42
|
+
## Your first five minutes
|
|
155
43
|
|
|
156
|
-
|
|
157
|
-
|
|
44
|
+
1. Open **Overview**. Token Harness detects your agents, installed optimizers and integration health.
|
|
45
|
+
2. **Review optimizer setup.** Start with RTK + HarnessTrim where the exact combination is supported.
|
|
46
|
+
Each optimizer has its setup, verification and removal controls together.
|
|
47
|
+
3. **Preview and apply.** Check the proposed changes, then approve. Opening the dashboard alone does
|
|
48
|
+
not change agent configuration.
|
|
49
|
+
4. **Authorize native hooks once.** In Codex, use `/hooks` to enable and trust the installed hooks.
|
|
50
|
+
In Claude Code, ensure hooks are enabled and start a fresh session after changing them.
|
|
51
|
+
5. **Use your agent normally.** Return to **Results** to inspect observed evidence. “Configured” means
|
|
52
|
+
setup exists; “runtime observed” means a qualifying callback or operation was actually recorded.
|
|
158
53
|
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
| [HarnessTrim](https://github.com/giuliastro/HarnessTrim) | Deterministic output/context reduction | Managed first-party integration |
|
|
163
|
-
| mcptoon | MCP discovery / compact manifest guidance | Optional managed integration on exact reviewed 0.7.10 rows; no savings assumed |
|
|
164
|
-
| GitNexus | Repository graph / MCP context | Optional managed Claude/Codex integration for reviewed 1.6.12; license review required |
|
|
165
|
-
| Headroom | Local MCP context compression/retrieval | Optional config-only managed Claude/Codex integration for already-installed 0.37.0; package prerequisite stays user-owned |
|
|
166
|
-
| [cclimits](https://github.com/cruzanstx/cclimits) | Optional Claude allowance evidence | Read-only evidence; not an optimizer |
|
|
167
|
-
| [ccusage](https://github.com/ccusage/ccusage) | Local usage history | Read-only evidence; never subscription quota |
|
|
168
|
-
|
|
169
|
-
Provider compatibility is deliberately **not pinned forever to the first fixture version**. The
|
|
170
|
-
current compatibility policy includes RTK **0.49.0** (source-contract reviewed; the latest live
|
|
171
|
-
Windows harness-mutation fixture is 0.48.0) and HarnessTrim **0.3.0**. Newer HarnessTrim builds can
|
|
172
|
-
be accepted without another hard-coded version bump when their executable version matches their
|
|
173
|
-
machine-readable `capabilities` version and the semantic surface/write-set comparison reports no
|
|
174
|
-
drift.
|
|
175
|
-
|
|
176
|
-
Token Harness and provider **package updates are separate from harness configuration writes**.
|
|
177
|
-
`token-harness update` can update the app when it is running from its verified global npm
|
|
178
|
-
installation, and can replace a reviewed provider target without requiring an exact historical
|
|
179
|
-
Claude/Codex fixture for that package version; exact compatibility rows still gate any later
|
|
180
|
-
managed agent-config mutation. HarnessTrim updates use its reviewed npm channel and capture the
|
|
181
|
-
previous global version for rollback. On native Windows RTK still prefers WinGet, but when that catalog is behind the
|
|
182
|
-
reviewed 0.49.0 target Token Harness can fall back to the exact official GitHub Windows x64 release:
|
|
183
|
-
it verifies GitHub's published SHA-256, replaces only the uniquely resolved `rtk.exe`, verifies the
|
|
184
|
-
new version, and restores and re-verifies the previous bytes on failure. This package-only fallback
|
|
185
|
-
does not widen RFC 0009 or grant permission to mutate agent configuration.
|
|
186
|
-
|
|
187
|
-
RTK has no equivalent machine-readable capability endpoint, so releases newer than the explicitly
|
|
188
|
-
reviewed RTK set remain visible as `unknown-newer` until their consumed contract is checked. See
|
|
189
|
-
[docs/provider-version-compatibility.md](docs/provider-version-compatibility.md).
|
|
190
|
-
|
|
191
|
-
**Codex hook activation is manual.** Token Harness can write an RTK hook declaration to
|
|
192
|
-
`hooks.json`, but Codex separately requires the hook to be enabled and trusted in Codex. After
|
|
193
|
-
setup, open Codex and enable/trust the hook. Token Harness never grants trust. A declaration in
|
|
194
|
-
`hooks.json` is config-only evidence until Codex reports the hook enabled and trusted; run
|
|
195
|
-
`token-harness verify --harness codex` to inspect the available verification tier.
|
|
196
|
-
|
|
197
|
-
Historical evaluation evidence remains available for mcptoon, GitNexus and Headroom. Detection or a promising
|
|
198
|
-
benchmark is not enough for a production-stack promotion or savings claim. Their campaign assessment is structured evidence for the
|
|
199
|
-
selection gate, not an activation or promotion decision. A candidate must pass structured promotion
|
|
200
|
-
readiness across benchmark capability, category fit, selection evidence, real activation
|
|
201
|
-
verification, managed lifecycle, compatibility/reversibility, project maturity and combined-stack
|
|
202
|
-
validation. Broader context owners also require an explicit admission decision.
|
|
203
|
-
|
|
204
|
-
See [docs/optimizer-priorities.md](docs/optimizer-priorities.md) and
|
|
205
|
-
[RFC 0027](docs/rfcs/0027-optimization-stack-manager.md).
|
|
206
|
-
|
|
207
|
-
## Stable-stack operating model
|
|
208
|
-
|
|
209
|
-
The intended lifecycle is:
|
|
210
|
-
|
|
211
|
-
```text
|
|
212
|
-
discover -> evaluate -> recommend -> install/configure -> verify -> measure
|
|
213
|
-
-> monitor -> update/re-evaluate -> rollback/uninstall
|
|
214
|
-
```
|
|
54
|
+
You do not need to keep the dashboard open for installed native hooks to run. Evidence appears after
|
|
55
|
+
qualifying activity; an idle installation cannot prove execution. Use the action beside an affected
|
|
56
|
+
agent or optimizer when a version or prerequisite needs attention.
|
|
215
57
|
|
|
216
|
-
|
|
217
|
-
agent/optimizer changes version, configuration drift appears, measured value deteriorates, quality
|
|
218
|
-
regresses, workload shape changes materially, or a credible better candidate appears.
|
|
58
|
+
## What you get
|
|
219
59
|
|
|
220
|
-
|
|
60
|
+
| Feature | How it helps |
|
|
61
|
+
| --- | --- |
|
|
62
|
+
| **One optimization dashboard** | Inspect agents, optimizers, health and updates together. |
|
|
63
|
+
| **Less noisy tool output** | Connect RTK and HarnessTrim on reviewed integration rows. |
|
|
64
|
+
| **Automatic prompt guidance** | Opt into native hooks that supply a bounded delegation policy on each prompt. |
|
|
65
|
+
| **Allowance-aware advice** | Inspect five-hour/weekly windows and native model/effort recommendations when evidence is available. |
|
|
66
|
+
| **Measured results** | Search and filter sources; expand each result for before/after values and attribution. |
|
|
67
|
+
| **Safe maintenance** | Preview changes, retain backups, verify updates and remove owned configuration. |
|
|
221
68
|
|
|
222
|
-
|
|
223
|
-
|
|
69
|
+
Automatic routing needs no skill invocation or prompt prefix after setup and native authorization.
|
|
70
|
+
The root model stays in place; the agent may delegate an eligible bounded subtask to a cheaper native
|
|
71
|
+
model and review the result. A callback proves the hook ran, not that a particular child model was
|
|
72
|
+
used or that allowance was saved. See [routing details](docs/rfcs/0030-automatic-native-prompt-routing.md).
|
|
224
73
|
|
|
225
|
-
|
|
226
|
-
example:
|
|
74
|
+
## Savings and statistics: what is measured today?
|
|
227
75
|
|
|
228
|
-
|
|
229
|
-
|
|
230
|
-
|
|
231
|
-
```
|
|
76
|
+
**There is no verified universal “save X%” claim.** Your results depend on the workload, installed
|
|
77
|
+
stack and available observations. The dashboard reports evidence from your own machine rather than
|
|
78
|
+
turning upstream marketing numbers into your savings.
|
|
232
79
|
|
|
233
|
-
|
|
234
|
-
|
|
80
|
+
| Measurement | What you can trust |
|
|
81
|
+
| --- | --- |
|
|
82
|
+
| **Output reduction** | Attributable provider receipts, with tokens/characters and exact/estimated classes labelled separately. This is not a subscription saving percentage. |
|
|
83
|
+
| **Subscription allowance** | Same-task baseline/optimized comparisons using authoritative five-hour and weekly observations, without crossing resets. Quality and retries are checked separately. |
|
|
84
|
+
| **Routing benefit** | Paired runs with routing disabled/enabled, genuine callback evidence and both quality gates passed. No qualifying pair means **Not measured yet**. |
|
|
85
|
+
| **API cost** | Attributable billed tokens and a verified model-price basis are required. Without them, cost stays **Not measured yet**. |
|
|
86
|
+
| **Quality** | Acceptance, retries and regressions remain visible. A regression blocks a positive saving claim. |
|
|
235
87
|
|
|
236
|
-
The
|
|
237
|
-
|
|
238
|
-
|
|
88
|
+
The reported mcptoon evaluation completed **8 paired tasks** and **16/16 quality passes**:
|
|
89
|
+
4 pairs were equivalent, 2 favoured the optimized variant and 2 favoured the baseline. It produced
|
|
90
|
+
**no consistent subscription saving signal**, and no successful mcptoon activity was observed inside
|
|
91
|
+
the optimized task windows. Those differences cannot be attributed to mcptoon. It remains optional
|
|
92
|
+
and experimental, rather than a proven third savings mechanism.
|
|
93
|
+
[Evaluation record](docs/development-status.md#historical-development-log).
|
|
239
94
|
|
|
240
|
-
|
|
241
|
-
|
|
242
|
-
prompt; you do not invoke the skill or prefix prompts with a Token Harness command. The root model
|
|
243
|
-
does not change. The agent may delegate one eligible bounded subtask to a native cheaper model,
|
|
244
|
-
then integrate and review its work. Codex asks for `gpt-6-luna`; Claude Code asks for its current
|
|
245
|
-
Haiku alias when available. Claude takes effect in a new session; Codex requires reviewing and
|
|
246
|
-
trusting the hook in `/hooks`. The dashboard separates configured hooks from callbacks observed at
|
|
247
|
-
runtime, and shows model-routing savings only after a paired quality-gated benchmark. Local token
|
|
248
|
-
counts and 5-hour/weekly allowance percentages remain separate measurements. See
|
|
249
|
-
[RFC 0030](docs/rfcs/0030-automatic-native-prompt-routing.md).
|
|
95
|
+
Releases are gated by CI on **Windows, macOS and Linux**, plus exact-artifact publication/install
|
|
96
|
+
checks. Real Windows combined-stack evidence and broader promotion gates remain open. [Release evidence and roadmap status](docs/plan-status.md).
|
|
250
97
|
|
|
251
|
-
|
|
252
|
-
target path and actions, then apply the returned plan id:
|
|
98
|
+
For a terminal summary of locally recorded evidence:
|
|
253
99
|
|
|
254
100
|
```sh
|
|
255
|
-
token-harness
|
|
256
|
-
token-harness apply --plan <plan-id> --yes
|
|
101
|
+
token-harness savings --since 7d
|
|
257
102
|
```
|
|
258
103
|
|
|
259
|
-
|
|
260
|
-
|
|
104
|
+
No provider totals are silently combined across incompatible units. Missing evidence is unknown,
|
|
105
|
+
not zero. Positive savings claims require attributable paired measurements with quality gates.
|
|
261
106
|
|
|
262
|
-
|
|
263
|
-
token-harness plan --harness claude --provider none --agent-routing --json
|
|
264
|
-
token-harness apply --plan <plan-id> --yes
|
|
265
|
-
```
|
|
266
|
-
|
|
267
|
-
## Advanced CLI
|
|
107
|
+
## Available optimizers
|
|
268
108
|
|
|
269
|
-
|
|
270
|
-
browser controller itself.
|
|
271
|
-
|
|
272
|
-
| Command | Purpose | Changes agent/project config? |
|
|
109
|
+
| Component | Purpose | Current role |
|
|
273
110
|
| --- | --- | --- |
|
|
274
|
-
|
|
|
275
|
-
|
|
|
276
|
-
|
|
|
277
|
-
|
|
|
278
|
-
|
|
|
279
|
-
|
|
|
280
|
-
| `apply` | Apply a reviewed stored plan | Yes, only with `--yes` |
|
|
281
|
-
| `verify` | Check the declared integration tier | No |
|
|
282
|
-
| `metrics` | Report attributable reducer savings | No |
|
|
283
|
-
| `status` | Report pipelines, drift and importer modes | No |
|
|
284
|
-
| `update` | Check/update reviewed provider packages | Yes, only with `--yes` |
|
|
285
|
-
| `rollback` | Restore the latest transaction snapshot | Yes, only with `--yes` |
|
|
286
|
-
| `uninstall` | Remove owned integration entries | Yes, only with `--yes` |
|
|
287
|
-
| `schedule` | Compare Claude Code and Codex using available evidence | No |
|
|
288
|
-
| `handoff` | Build a bounded cross-agent handoff | No |
|
|
289
|
-
| `benchmark*`, `transfer*` | Capture and compare empirical evidence | Local state only |
|
|
290
|
-
|
|
291
|
-
Need stable machine-readable output? Add `--json`. Need the evidence behind a human summary? Add
|
|
292
|
-
`--verbose`.
|
|
293
|
-
|
|
294
|
-
The older automation contracts remain available. `ui --json` preserves its existing schema-1
|
|
295
|
-
report; `ui --read-only` opens the legacy read-only UI; `ui --no-open` starts the guided app without
|
|
296
|
-
launching a browser.
|
|
297
|
-
|
|
298
|
-
### Evaluation evidence (advanced / maintainers)
|
|
299
|
-
|
|
300
|
-
Evaluation campaigns are an advanced maintainer workflow; managed setup stays in the unified **Optimization Stack**. The app
|
|
301
|
-
keeps a resumable campaign ID for each candidate/harness pair and reads campaign progress, assessment
|
|
302
|
-
and the exact **Next** step directly in the browser. Use **Start baseline capture** or **Start optimized
|
|
303
|
-
capture**, run the requested task in the selected coding agent, then choose **Record outcome** and
|
|
304
|
-
enter the quality/attempt values you actually observed. Normal use no longer requires copying
|
|
305
|
-
`benchmark-start` or `benchmark-finish` commands into a terminal.
|
|
306
|
-
|
|
307
|
-
The equivalent advanced CLI flow starts by asking the campaign engine for its current state. This
|
|
308
|
-
GitNexus example intentionally uses the only currently reviewed campaign row:
|
|
309
|
-
|
|
310
|
-
```sh
|
|
311
|
-
token-harness benchmark-matrix \
|
|
312
|
-
--benchmark-id gitnexus-claude-eval-1 \
|
|
313
|
-
--candidate gitnexus \
|
|
314
|
-
--harness claude
|
|
315
|
-
```
|
|
316
|
-
|
|
317
|
-
Follow only the **Next** command printed by that report, complete the task honestly, then rerun the
|
|
318
|
-
same `benchmark-matrix` command. Before an optimized run, enable the candidate through its own
|
|
319
|
-
documented workflow. Token Harness records the experiment target but does not treat attribution—or
|
|
320
|
-
the browser acknowledgement—as proof that the candidate was active. For GitNexus on the reviewed
|
|
321
|
-
Claude Code `2.1.269` × GitNexus `1.6.12` × native-Linux row, the harness-native MCP inventory can
|
|
322
|
-
prove only that the GitNexus server was available at both task boundaries. That is not proof Claude
|
|
323
|
-
actually called a GitNexus tool, so the activation-verification promotion gate remains blocked
|
|
324
|
-
without a separate reviewed usage witness. See
|
|
325
|
-
[`docs/candidates/gitnexus-real-campaign.md`](docs/candidates/gitnexus-real-campaign.md).
|
|
326
|
-
|
|
327
|
-
The selection assessment can become decision-ready after enough evidence across task classes, but it
|
|
328
|
-
still cannot promote a candidate by itself. The remaining lifecycle and combined-stack gates must be
|
|
329
|
-
satisfied separately.
|
|
330
|
-
|
|
331
|
-
### Workload-aware allowance planning
|
|
332
|
-
|
|
333
|
-
If you explicitly know the remaining backlog, the advanced CLI can reason about whether that work
|
|
334
|
-
fits the currently observed allowance:
|
|
111
|
+
| [RTK](https://github.com/rtk-ai/rtk) | Reduce shell/tool output | Baseline integration on reviewed rows. |
|
|
112
|
+
| [HarnessTrim](https://github.com/giuliastro/HarnessTrim) | Deterministic output/context reduction | Baseline integration on reviewed rows. |
|
|
113
|
+
| mcptoon | Compact MCP discovery/manifests | Optional; install through an existing pipx or uv. Missing prerequisites show installation options. |
|
|
114
|
+
| GitNexus | Repository graph and MCP context | Optional reviewed integration; package/index preparation and license review are user responsibilities. |
|
|
115
|
+
| Headroom | MCP context compression/retrieval | Optional configuration-only integration for an already-installed reviewed package. |
|
|
116
|
+
| [cclimits](https://github.com/cruzanstx/cclimits) / [ccusage](https://github.com/ccusage/ccusage) | Allowance / usage observations | Read-only evidence sources, rather than optimizers. Usage history is not remaining quota. |
|
|
335
117
|
|
|
336
|
-
|
|
337
|
-
|
|
338
|
-
|
|
339
|
-
|
|
340
|
-
|
|
341
|
-
For a mixed queued workload:
|
|
342
|
-
|
|
343
|
-
```sh
|
|
344
|
-
token-harness schedule --current codex --candidate claude \
|
|
345
|
-
--workload mechanical=2,standard=3,hard=1
|
|
346
|
-
```
|
|
347
|
-
|
|
348
|
-
Token Harness does not infer remaining tasks from session length, local tokens or raw provider
|
|
349
|
-
percentages. If the required benchmark/allowance evidence is incomplete, capacity remains unknown.
|
|
350
|
-
See [RFC 0020](docs/rfcs/0020-workload-aware-allowance.md) and
|
|
351
|
-
[RFC 0021](docs/rfcs/0021-mixed-workload-allocation.md).
|
|
352
|
-
|
|
353
|
-
### Applying native recommendations from the CLI
|
|
354
|
-
|
|
355
|
-
`optimize` remains read-only. The explicit CLI path is review then apply:
|
|
356
|
-
|
|
357
|
-
```sh
|
|
358
|
-
token-harness plan --harness claude --native-policy --task mechanical --profile economy
|
|
359
|
-
token-harness apply --plan <printed-plan-id> --yes
|
|
360
|
-
```
|
|
118
|
+
Current source-reviewed package targets are **RTK 0.51.0** and **HarnessTrim 0.3.1**. Package update
|
|
119
|
+
policy and exact agent-configuration compatibility are separate: a newer installed package does not
|
|
120
|
+
automatically admit new config writes or a combined-stack review. RTK + HarnessTrim still need real
|
|
121
|
+
combined workload evidence before broader promotion. Optional integrations do not count as proven
|
|
122
|
+
savings mechanisms merely because they are installed.
|
|
361
123
|
|
|
362
|
-
|
|
124
|
+
See [version policy](docs/provider-version-compatibility.md),
|
|
125
|
+
[compatibility and verification tiers](docs/matrices.md), and
|
|
126
|
+
[candidate promotion gates](docs/candidates/promotion-readiness.md).
|
|
363
127
|
|
|
364
|
-
##
|
|
128
|
+
## Update, disconnect or undo
|
|
365
129
|
|
|
366
|
-
|
|
130
|
+
The dashboard checks for updates on opening. When an update is available, review its exact version
|
|
131
|
+
and choose **Install updates** to approve installation. It re-checks health automatically afterward.
|
|
132
|
+
A supported application update offers **Restart and re-check** to load the new version; if replacement startup fails, the
|
|
133
|
+
current dashboard stays available.
|
|
367
134
|
|
|
368
|
-
|
|
369
|
-
- a browser configuration mutation requires preview and explicit approval;
|
|
370
|
-
- guided candidate capture buttons write only bounded local benchmark state, are CSRF-protected and
|
|
371
|
-
must still match the campaign engine's current step immediately before the write;
|
|
372
|
-
- a browser activation acknowledgement is never treated as verified candidate activation;
|
|
373
|
-
- CLI mutations require their explicit `--yes` form;
|
|
374
|
-
- plans are checked again immediately before apply;
|
|
375
|
-
- existing files are backed up before a managed write;
|
|
376
|
-
- only exact Token Harness-owned entries are removed by uninstall;
|
|
377
|
-
- newer or untested combinations are reported rather than guessed;
|
|
378
|
-
- an available provider update outside reviewed package compatibility is kept out rather than forced;
|
|
379
|
-
- provider package replacement does not bypass the stricter compatibility gate for harness config writes;
|
|
380
|
-
- the guided app binds only to `127.0.0.1` and protects local controls with Host/Origin checks, a
|
|
381
|
-
per-process anti-forgery token and single-use approval tickets;
|
|
382
|
-
- source code, prompts, command contents, credentials and cookies are not sent to a Token Harness
|
|
383
|
-
service.
|
|
384
|
-
|
|
385
|
-
Plans, receipts, metrics and backups stay in the local Token Harness state directory.
|
|
386
|
-
|
|
387
|
-
See [RFC 0013](docs/rfcs/0013-guided-local-experience.md),
|
|
388
|
-
[RFC 0004](docs/rfcs/0004-safety-and-installation.md), and
|
|
389
|
-
[RFC 0006](docs/rfcs/0006-cli-contract.md).
|
|
390
|
-
|
|
391
|
-
## Updating, checking and undoing
|
|
392
|
-
|
|
393
|
-
Update Token Harness itself:
|
|
135
|
+
To update manually:
|
|
394
136
|
|
|
395
137
|
```sh
|
|
396
138
|
npm install --global token-harness@latest
|
|
139
|
+
token-harness
|
|
397
140
|
```
|
|
398
141
|
|
|
399
|
-
|
|
400
|
-
|
|
401
|
-
it offers **Install updates** in the same dialog. After updating Token Harness from the browser,
|
|
402
|
-
restart the app to load the new version. From the advanced CLI, the preview prints the exact
|
|
403
|
-
confirmation command:
|
|
404
|
-
|
|
405
|
-
```sh
|
|
406
|
-
token-harness update
|
|
407
|
-
token-harness update --yes
|
|
408
|
-
```
|
|
409
|
-
|
|
410
|
-
`update` installs Token Harness itself only when the running copy matches its global npm package;
|
|
411
|
-
other install methods remain available through their original updater. Optimizer updates replace
|
|
412
|
-
only installed providers whose target is inside the reviewed provider-package policy. The npm
|
|
413
|
-
package inventory captures the previous Token Harness version for rollback. HarnessTrim uses npm
|
|
414
|
-
and captures the previous global version for rollback. On native
|
|
415
|
-
Windows RTK prefers WinGet; when WinGet cannot yet reach the reviewed target, Token Harness can use
|
|
416
|
-
the verified official GitHub Windows x64 release fallback described above. That fallback verifies the
|
|
417
|
-
published digest and post-update version and restores the previous executable on failure.
|
|
418
|
-
|
|
419
|
-
Remove only Token Harness-owned integration entries:
|
|
142
|
+
To disconnect an optimizer, use **Remove managed setup** beside that optimizer.
|
|
143
|
+
It removes only entries Token Harness owns. The CLI equivalent for owned integrations is:
|
|
420
144
|
|
|
421
145
|
```sh
|
|
422
146
|
token-harness uninstall --yes
|
|
423
147
|
```
|
|
424
148
|
|
|
425
|
-
|
|
149
|
+
To restore the latest complete configuration snapshot:
|
|
426
150
|
|
|
427
151
|
```sh
|
|
428
152
|
token-harness rollback --yes
|
|
429
153
|
```
|
|
430
154
|
|
|
431
|
-
|
|
432
|
-
|
|
155
|
+
Rollback restores whole files and can revert later manual edits. Prefer owned removal when you only
|
|
156
|
+
want to disconnect Token Harness integrations. Provider package management and uninstall scope are
|
|
157
|
+
shown in the reviewed plan.
|
|
433
158
|
|
|
434
|
-
##
|
|
159
|
+
## Local by design
|
|
435
160
|
|
|
436
|
-
|
|
161
|
+
The dashboard binds to `127.0.0.1`. Plans, receipts, metrics and backups stay in local state.
|
|
162
|
+
Token Harness does not send your source code, prompts, command contents or credentials to a Token
|
|
163
|
+
Harness service. Your coding agent and any explicitly enabled provider keep their own network
|
|
164
|
+
behaviour.
|
|
437
165
|
|
|
438
|
-
|
|
439
|
-
|
|
440
|
-
unsupported
|
|
166
|
+
A configuration change requires an explicit review and approval; plans are checked again before apply.
|
|
167
|
+
Managed writes retain backups, preserve unrelated configuration and support verification and rollback. Unknown hook formats
|
|
168
|
+
or unsupported configuration rows require evidence instead of guessed writes.
|
|
441
169
|
|
|
442
|
-
|
|
170
|
+
## Need help?
|
|
443
171
|
|
|
444
|
-
|
|
445
|
-
|
|
446
|
-
|
|
447
|
-
|
|
448
|
-
|
|
172
|
+
- **Command not found:** check `node --version` and `npm list --global token-harness`, then reopen
|
|
173
|
+
the terminal. Node must be at least 22.13.
|
|
174
|
+
- **Routing configured but no callback:** check the hook's enabled/trusted state in `/hooks`, start
|
|
175
|
+
a new session and submit an ordinary prompt. You should never need to invoke the skill per prompt.
|
|
176
|
+
- **Optimizer unavailable:** open its setup/installation options; some dependencies remain
|
|
177
|
+
user-installed prerequisites and some configuration rows have narrower support.
|
|
178
|
+
- **Quota unavailable:** a missing observation is not zero allowance. Use the status explanation;
|
|
179
|
+
do not paste credentials to troubleshoot it.
|
|
180
|
+
- **Integration needs attention:** inspect `token-harness doctor --verbose` and
|
|
181
|
+
`token-harness verify --verbose`; the reported verification tier explains what was actually checked.
|
|
449
182
|
|
|
450
|
-
|
|
183
|
+
[Windows notes](docs/windows-companions.md) · [CLI and evaluation guide](docs/cli-guide.md) ·
|
|
184
|
+
[Agent Skill](skills/token-harness/SKILL.md) · [Latest release](https://github.com/giuliastro/token-harness/releases/latest)
|
|
451
185
|
|
|
452
|
-
|
|
186
|
+
## Contribute and follow the plan
|
|
453
187
|
|
|
454
|
-
|
|
455
|
-
|
|
456
|
-
npm list --global token-harness
|
|
457
|
-
```
|
|
458
|
-
|
|
459
|
-
Node must be at least 22.13. Reopen the terminal after installation if needed.
|
|
460
|
-
|
|
461
|
-
### Setup or verification needs attention
|
|
462
|
-
|
|
463
|
-
Use the action shown beside the affected agent or optimizer in Overview. For technical evidence:
|
|
464
|
-
|
|
465
|
-
```sh
|
|
466
|
-
token-harness doctor --verbose
|
|
467
|
-
token-harness verify --verbose
|
|
468
|
-
```
|
|
469
|
-
|
|
470
|
-
Do not force an unsupported plan. A newer version outside reviewed compatibility is normally a
|
|
471
|
-
safety limitation, not a reason to overwrite the known-working installation.
|
|
188
|
+
[Current progress and next steps](docs/plan-status.md) · [Full plan](PLAN.md) ·
|
|
189
|
+
[Accepted RFCs](docs/rfcs) · [Release readiness](docs/release-readiness.md)
|
|
472
190
|
|
|
473
|
-
|
|
474
|
-
|
|
475
|
-
From an existing clone:
|
|
476
|
-
|
|
477
|
-
```sh
|
|
478
|
-
npm start
|
|
479
|
-
```
|
|
480
|
-
|
|
481
|
-
For a fresh clone:
|
|
191
|
+
To run a fresh source checkout:
|
|
482
192
|
|
|
483
193
|
```sh
|
|
484
194
|
git clone https://github.com/giuliastro/token-harness.git
|
|
@@ -487,35 +197,14 @@ npx --yes pnpm@10.33.4 install --frozen-lockfile
|
|
|
487
197
|
npm start
|
|
488
198
|
```
|
|
489
199
|
|
|
490
|
-
|
|
491
|
-
|
|
492
|
-
|
|
493
|
-
## Development
|
|
494
|
-
|
|
495
|
-
```sh
|
|
496
|
-
git clone https://github.com/giuliastro/token-harness.git
|
|
497
|
-
cd token-harness
|
|
498
|
-
npx --yes pnpm@10.33.4 install --frozen-lockfile
|
|
499
|
-
npx --yes pnpm@10.33.4 typecheck
|
|
500
|
-
npx --yes pnpm@10.33.4 lint
|
|
501
|
-
npx --yes pnpm@10.33.4 test
|
|
502
|
-
npx --yes pnpm@10.33.4 build
|
|
503
|
-
npx --yes pnpm@10.33.4 smoke
|
|
504
|
-
npx --yes pnpm@10.33.4 package
|
|
505
|
-
npx --yes pnpm@10.33.4 smoke:install
|
|
506
|
-
```
|
|
507
|
-
|
|
508
|
-
Using `corepack enable` is optional. On a system-wide Windows Node installation it can require
|
|
509
|
-
administrator permission to modify `C:\Program Files\nodejs`; the `npx pnpm@10.33.4` form above
|
|
510
|
-
does not require that Corepack shim write.
|
|
200
|
+
For an existing clone with dependencies installed, `npm start` builds and opens the local source.
|
|
201
|
+
Source changes do not reach npm users until a release is published.
|
|
511
202
|
|
|
512
|
-
|
|
513
|
-
|
|
514
|
-
|
|
515
|
-
[
|
|
516
|
-
[RFCs](docs/rfcs).
|
|
203
|
+
Development checks: `pnpm typecheck`, `pnpm lint`, `pnpm test`, `pnpm build`, `pnpm smoke`,
|
|
204
|
+
`pnpm package`, `pnpm smoke:install`. Use the pinned pnpm version above if Corepack is unavailable;
|
|
205
|
+
`corepack enable` is optional and can require administrator rights on Windows.
|
|
206
|
+
Read [contributor instructions](AGENTS.md) and accepted RFCs before changing public contracts.
|
|
517
207
|
|
|
518
208
|
## License
|
|
519
209
|
|
|
520
|
-
[Apache License 2.0](LICENSE). Referenced
|
|
521
|
-
licenses.
|
|
210
|
+
[Apache License 2.0](LICENSE). Referenced tools are independent projects with their own licenses.
|