token-harness 0.1.10 → 0.1.11
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +370 -326
- package/package.json +1 -1
- package/sbom.json +3 -3
- package/token-harness.mjs +13623 -5701
package/README.md
CHANGED
|
@@ -1,480 +1,524 @@
|
|
|
1
1
|
# Token Harness
|
|
2
2
|
|
|
3
|
-
**
|
|
3
|
+
**Build, verify and measure an optimization stack for Claude Code and Codex.**
|
|
4
4
|
|
|
5
|
-
Token Harness checks your coding agents,
|
|
6
|
-
|
|
7
|
-
|
|
5
|
+
Token Harness is a local **optimization stack manager**. It checks your coding agents, manages the
|
|
6
|
+
optimization components it can safely own, keeps experimental candidates separate, verifies the
|
|
7
|
+
result, and reports savings only when it has evidence to support them.
|
|
8
8
|
|
|
9
|
-
|
|
9
|
+
It is not another coding agent and it does not replace specialized projects such as RTK or
|
|
10
|
+
HarnessTrim.
|
|
10
11
|
|
|
11
|
-
|
|
12
|
-
|
|
12
|
+
## Start here
|
|
13
|
+
|
|
14
|
+
Requirements:
|
|
15
|
+
|
|
16
|
+
- Node.js 22.13 or newer;
|
|
17
|
+
- Claude Code or Codex installed;
|
|
18
|
+
- the coding agent you want to use already signed in.
|
|
19
|
+
|
|
20
|
+
Install and open Token Harness:
|
|
13
21
|
|
|
14
22
|
```sh
|
|
15
23
|
npm install --global token-harness@latest
|
|
16
24
|
token-harness
|
|
17
25
|
```
|
|
18
26
|
|
|
19
|
-
The browser is
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
27
|
+
That is the normal human workflow. The browser app is the primary interface; there is no daily list
|
|
28
|
+
of CLI commands to memorize.
|
|
29
|
+
|
|
30
|
+
### First run
|
|
31
|
+
|
|
32
|
+
1. Open **Dashboard** and let Token Harness inspect the current setup.
|
|
33
|
+
2. If setup is incomplete, choose **Open setup**.
|
|
34
|
+
3. In **Setup**, work from top to bottom:
|
|
35
|
+
- Coding agents
|
|
36
|
+
- Managed optimizers
|
|
37
|
+
- Experimental tools
|
|
38
|
+
- Optional agent tuning
|
|
39
|
+
- Checks and maintenance
|
|
40
|
+
4. For Claude Code or Codex, choose **Review setup** when a managed setup is available.
|
|
41
|
+
5. Read the exact proposed changes. **Apply reviewed setup** appears only when there is a concrete
|
|
42
|
+
safe plan to apply.
|
|
43
|
+
6. Keep using Claude Code or Codex normally.
|
|
44
|
+
7. Open **Results** when you want to see what Token Harness can actually prove.
|
|
45
|
+
|
|
46
|
+
Opening the app does not change your configuration. Read-only checks stay read-only, and a managed
|
|
47
|
+
write requires an explicit review and approval.
|
|
23
48
|
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
49
|
+
## The three views
|
|
50
|
+
|
|
51
|
+
### Dashboard
|
|
52
|
+
|
|
53
|
+
Dashboard answers the questions that matter first:
|
|
54
|
+
|
|
55
|
+
- is my setup ready;
|
|
56
|
+
- which coding agents and managed optimizers are active;
|
|
57
|
+
- what should I do next;
|
|
58
|
+
- what value has actually been measured;
|
|
59
|
+
- whether quality or an integration needs attention.
|
|
28
60
|
|
|
29
|
-
|
|
61
|
+
The headline cards deliberately distinguish measured evidence from unknown values. Missing evidence
|
|
62
|
+
is never shown as zero savings.
|
|
30
63
|
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
imports available provider records and refreshes its readings automatically.
|
|
64
|
+
### Setup
|
|
65
|
+
|
|
66
|
+
Setup is one ordered workflow instead of a collection of unrelated actions.
|
|
35
67
|
|
|
36
|
-
**
|
|
37
|
-
provider, measurement class, units, changed-output counts, and before/after values. It does not
|
|
38
|
-
add incompatible provider figures together. Negative results remain visible. No telemetry is
|
|
39
|
-
shown as **not measured**, never a reassuring zero or an invented subscription saving.
|
|
40
|
-
Some provider records may predate Token Harness; locally stored records are not guaranteed
|
|
41
|
-
complete lifetime history. RTK history is imported directly. HarnessTrim project-local records
|
|
42
|
-
must have been imported from their project, or exposed through a configured known metrics path;
|
|
43
|
-
the app does not crawl your disk looking for private projects.
|
|
68
|
+
**1. Coding agents**
|
|
44
69
|
|
|
45
|
-
|
|
70
|
+
Token Harness currently supports guided setup for Claude Code and Codex. Agent details also show
|
|
71
|
+
useful read-only allowance and connected-tool observations when available.
|
|
46
72
|
|
|
47
|
-
|
|
48
|
-
token-harness savings
|
|
49
|
-
```
|
|
73
|
+
**2. Managed optimizers**
|
|
50
74
|
|
|
51
|
-
|
|
52
|
-
|
|
75
|
+
The managed stack currently consists of **RTK + HarnessTrim** on individually reviewed
|
|
76
|
+
combinations. Token Harness tracks the exact combined provider set separately: if no combined-stack
|
|
77
|
+
review is recorded, Setup says so and keeps the stack incomplete rather than inferring compatibility
|
|
78
|
+
from healthy individual checks. Token Harness can prepare their integration transactionally, show the
|
|
79
|
+
exact plan, apply it only after approval, verify it, and remove only configuration it owns.
|
|
53
80
|
|
|
54
|
-
|
|
81
|
+
For maintainers validating the combined stack, `token-harness stack-review --json` captures the exact
|
|
82
|
+
configured provider versions and managed harness sets and reuses the existing passive `verify`
|
|
83
|
+
evidence for exact provider/harness pairs. Runtime evidence is credited only when it can be
|
|
84
|
+
attributed to that harness: HarnessTrim uses its native event harness field, while provider-wide
|
|
85
|
+
telemetry can be attributed by exclusion only when one harness is wired. With multiple harnesses, an
|
|
86
|
+
unattributable receipt stays **Unavailable** instead of being copied across rows. A `not-exercised`
|
|
87
|
+
result does not promise that ordinary agent use will create a receipt: the provider must record a
|
|
88
|
+
qualifying operation, and for reducers that means a real reduction. `stack-review` does not run an
|
|
89
|
+
active canary or spend a model call, and it never makes the compatibility decision itself. The
|
|
90
|
+
shipped combined-review registry stays empty until the captured configuration has enough real
|
|
91
|
+
verification evidence, has been benchmarked together, and has been deliberately reviewed; see
|
|
92
|
+
`docs/combined-stack-reviews.md`.
|
|
55
93
|
|
|
56
|
-
|
|
57
|
-
for one agent's rules at a time, and **Activity** for checks and guarded undo. Theme follows
|
|
58
|
-
your system; the header also offers light and dark modes.
|
|
94
|
+
**3. Experimental tools**
|
|
59
95
|
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
96
|
+
Headroom, mcptoon and GitNexus are visible as candidates, not silently promoted dependencies. Their
|
|
97
|
+
cards distinguish CLI installation, candidate-side activation/evaluation, Token Harness benchmark
|
|
98
|
+
evidence and promotion-readiness gates. External install or activation commands are shown for
|
|
99
|
+
review; Token Harness does not silently execute package managers, activate wrappers, index
|
|
100
|
+
repositories or register MCP servers.
|
|
64
101
|
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
102
|
+
Choose **Run standard evaluation** to start or resume a paired candidate campaign. Campaign state is
|
|
103
|
+
scoped to both the candidate and the selected harness, so Claude Code and Codex evidence cannot be
|
|
104
|
+
mixed accidentally. The browser reads the campaign directly and shows **Progress**, the current
|
|
105
|
+
selection signal, whether the evidence is **Decision ready**, the number of evidence-bearing pairs
|
|
106
|
+
and the exact **Next** step.
|
|
69
107
|
|
|
70
|
-
|
|
108
|
+
For normal use, the browser can now start and finish the local benchmark capture itself. You still
|
|
109
|
+
run the actual task in Claude Code or Codex. When the task finishes, record the quality result,
|
|
110
|
+
attempt count and failed-attempt count you actually observed. Before an optimized capture, Token
|
|
111
|
+
Harness requires you to acknowledge that you enabled the candidate through its own documented
|
|
112
|
+
workflow. That acknowledgement is **not activation verification** and is never treated as promotion
|
|
113
|
+
evidence. Every browser capture action is matched against the campaign engine's current step before
|
|
114
|
+
it can write local benchmark state, so a stale tab cannot advance a different step.
|
|
71
115
|
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
116
|
+
`benchmark-matrix`, `benchmark-start` and `benchmark-finish` remain available as advanced terminal
|
|
117
|
+
fallbacks for debugging or automation. The browser does not run the coding task, install or activate
|
|
118
|
+
a candidate, or claim that candidate attribution proves activation.
|
|
75
119
|
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
120
|
+
A campaign selection assessment can report `insufficient-evidence`, `promising`, `mixed` or
|
|
121
|
+
`negative`, plus whether the evidence is decision-ready. **Decision-ready is not promotion-ready.**
|
|
122
|
+
The campaign shows the same promotion-review gate count and next gate as the candidate card, using
|
|
123
|
+
the current local observation rather than starting another environment scan. Activation
|
|
124
|
+
verification, managed lifecycle, compatibility/reversibility, project maturity and combined-stack
|
|
125
|
+
validation remain separate gates.
|
|
80
126
|
|
|
81
|
-
**
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
127
|
+
Choose **Compare evaluation evidence** when you want one read-only view of campaigns you already
|
|
128
|
+
started. The comparison is loaded only on request, never creates a campaign, keeps candidate and
|
|
129
|
+
harness ordering fixed, and reports progress, selection signal, decision readiness, runtime
|
|
130
|
+
activation evidence and promotion gates without producing a composite score or automatic winner.
|
|
131
|
+
Activation is shown separately as **Verified**, **Blocked** or **Unreviewed**, with
|
|
132
|
+
verified/blocked/unknown pair counts when campaign evidence exists.
|
|
86
133
|
|
|
87
|
-
**
|
|
88
|
-
**Undo last change**, available after an application in that dashboard session, previews a
|
|
89
|
-
whole-file backup restoration. It refuses to undo a newer unrelated transaction. It restores
|
|
90
|
-
only the last successful agent transaction; manual edits to those same files after that
|
|
91
|
-
transaction would also be restored, as the confirmation explains.
|
|
134
|
+
**4. Optional agent tuning**
|
|
92
135
|
|
|
93
|
-
|
|
136
|
+
Reasoning preferences are separate from optimizer installation. They are persistent agent settings,
|
|
137
|
+
not hidden per-task switches, and are changed only through the normal preview/apply flow.
|
|
94
138
|
|
|
95
|
-
|
|
139
|
+
**5. Checks and maintenance**
|
|
96
140
|
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
```
|
|
141
|
+
Read-only integration checks, update checks and safe removal/undo controls live here. An update
|
|
142
|
+
outside reviewed compatibility is not forced.
|
|
100
143
|
|
|
101
|
-
|
|
144
|
+
### Results
|
|
102
145
|
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
cd token-harness
|
|
106
|
-
npx --yes pnpm@10.33.4 install --frozen-lockfile
|
|
107
|
-
npm start
|
|
108
|
-
```
|
|
146
|
+
Results keeps evidence separate from estimates. Depending on what is actually observable, it can
|
|
147
|
+
show:
|
|
109
148
|
|
|
110
|
-
|
|
111
|
-
|
|
149
|
+
- recorded optimizer output reduction;
|
|
150
|
+
- authoritative paired 5-hour / 7-day allowance evidence;
|
|
151
|
+
- API cost only when billed-token evidence and a verified price basis exist;
|
|
152
|
+
- paired quality evidence;
|
|
153
|
+
- experimental candidate evidence and campaign selection assessment when available;
|
|
154
|
+
- recent checks and changes from the current local app session.
|
|
112
155
|
|
|
113
|
-
|
|
156
|
+
**Not measured** means exactly that. Token Harness does not turn local token estimates into fake
|
|
157
|
+
subscription minutes, money or quota savings.
|
|
114
158
|
|
|
115
|
-
|
|
116
|
-
used by the app, automation, and optionally by a coding harness. Humans can keep using the browser
|
|
117
|
-
and the two entry points above.
|
|
159
|
+
## Daily use
|
|
118
160
|
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
skill mechanism can load it on demand, after which you can simply ask the harness to **use Token
|
|
122
|
-
Harness for this task**. The skill classifies substantial work conservatively, calls the existing
|
|
123
|
-
local `--json` optimizer at meaningful task boundaries, and can use explicit workload scheduling
|
|
124
|
-
when you have actually supplied a backlog. It does not run Token Harness before every tool call.
|
|
161
|
+
Keep launching `claude` or `codex` as usual. Deterministic optimizers that are installed, verified
|
|
162
|
+
and still beneficial are intended to remain enabled.
|
|
125
163
|
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
Local tokens are still not subscription quota, and raw Claude/Codex percentages are still not a
|
|
129
|
-
common currency.
|
|
164
|
+
Token Harness does **not** need to stay open, does not need a permanent background daemon, and does
|
|
165
|
+
not need to decide before every command whether an optimizer should run.
|
|
130
166
|
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
and wait for your explicit approval before `apply`. A saved preference may affect future sessions;
|
|
134
|
-
it is not silently presented as a live change to the current session.
|
|
167
|
+
Open `token-harness` when you want to inspect health/results, review a setup change, check an update,
|
|
168
|
+
verify integrations or re-evaluate the stack after a meaningful version/configuration change.
|
|
135
169
|
|
|
136
|
-
The
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
**Enabled externally** for a byte-identical user-owned skill, or an explicit not-enabled,
|
|
141
|
-
custom/conflict, or unavailable state. Existing `token-harness` skill directories are never
|
|
142
|
-
overwritten or silently adopted, and matching bytes alone never create an ownership claim. The
|
|
143
|
-
browser remains fully usable without any skill or second AI subscription. See
|
|
144
|
-
[RFC 0023](docs/rfcs/0023-guided-agent-skill-install.md) for the install and ownership boundary.
|
|
170
|
+
The app does not periodically reload the whole setup. A full read happens on initial open or when you
|
|
171
|
+
choose **Refresh**. Existing readings remain visible while a refresh runs. Applying a reviewed change
|
|
172
|
+
marks the displayed data as previous state instead of immediately launching another expensive full
|
|
173
|
+
read.
|
|
145
174
|
|
|
146
|
-
|
|
147
|
-
`verify`, `metrics`, `rollback`, and their JSON reports. `ui --json` preserves its existing
|
|
148
|
-
schema-1 report; `ui --read-only` opens the legacy read-only dashboard. `ui --no-open` starts
|
|
149
|
-
the guided app without launching a browser. Stop either local server with Ctrl+C. See
|
|
150
|
-
[RFC 0022](docs/rfcs/0022-agent-native-skill.md) for the agent-facing safety boundary.
|
|
175
|
+
## What counts as savings
|
|
151
176
|
|
|
152
|
-
|
|
177
|
+
Token Harness keeps different evidence classes separate.
|
|
153
178
|
|
|
154
|
-
|
|
179
|
+
**Recorded output savings** are attributable reducer measurements. Providers, units and measurement
|
|
180
|
+
classes are not silently added together. Negative results and errors remain visible.
|
|
155
181
|
|
|
156
|
-
|
|
157
|
-
|
|
182
|
+
**5h / 7d allowance savings** require authoritative paired before/after allowance evidence. A
|
|
183
|
+
five-hour percentage may also be expressed as the equivalent share of that 300-minute allowance
|
|
184
|
+
window. Weekly quota is not converted into seven days of wall-clock compute.
|
|
158
185
|
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
HarnessTrim: active on Codex
|
|
186
|
+
**API cost** stays **Not measured yet** until attributable billed input/output tokens and a verified
|
|
187
|
+
model-price basis are available.
|
|
162
188
|
|
|
163
|
-
|
|
164
|
-
|
|
189
|
+
**Quality** is measured independently. A measured regression blocks a positive allowance-saving
|
|
190
|
+
claim rather than letting a smaller token number win by itself.
|
|
165
191
|
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
```
|
|
169
|
-
|
|
170
|
-
A newer-than-tested combination is not presented as if the whole setup were broken:
|
|
171
|
-
|
|
172
|
-
```text
|
|
173
|
-
TOKEN HARNESS - READY WITH LIMITATIONS
|
|
192
|
+
Upstream benchmark numbers are useful for deciding what to test; they are never copied directly into
|
|
193
|
+
your savings total.
|
|
174
194
|
|
|
175
|
-
|
|
176
|
-
Claude Code: configured
|
|
177
|
-
RTK: active on Claude Code
|
|
195
|
+
For a terminal-only savings summary:
|
|
178
196
|
|
|
179
|
-
|
|
180
|
-
|
|
181
|
-
You can keep working; verify the active integrations when convenient.
|
|
197
|
+
```sh
|
|
198
|
+
token-harness savings
|
|
182
199
|
```
|
|
183
200
|
|
|
184
|
-
|
|
201
|
+
Optional windows are `--since 7d` and `--since 30d`.
|
|
185
202
|
|
|
186
|
-
|
|
187
|
-
token-harness doctor --verbose
|
|
188
|
-
```
|
|
203
|
+
## Current optimization stack
|
|
189
204
|
|
|
190
|
-
|
|
205
|
+
Token Harness prefers thin integrations around strong specialized projects instead of copying their
|
|
206
|
+
algorithms into this repository.
|
|
191
207
|
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
|
|
208
|
+
| Component | Role | Management |
|
|
209
|
+
| --- | --- | --- |
|
|
210
|
+
| [RTK](https://github.com/rtk-ai/rtk) | Shell/tool output reduction | Managed on reviewed combinations |
|
|
211
|
+
| [HarnessTrim](https://github.com/giuliastro/HarnessTrim) | Deterministic output/context reduction | Managed first-party integration |
|
|
212
|
+
| [cclimits](https://github.com/cruzanstx/cclimits) | Optional Claude allowance evidence | Read-only evidence; not an optimizer |
|
|
213
|
+
| [ccusage](https://github.com/ccusage/ccusage) | Local usage history | Read-only evidence; never subscription quota |
|
|
214
|
+
|
|
215
|
+
Provider compatibility is deliberately **not pinned forever to the first fixture version**. The
|
|
216
|
+
current compatibility policy includes RTK **0.49.0** (source-contract reviewed; the latest live
|
|
217
|
+
Windows harness-mutation fixture is 0.48.0) and HarnessTrim **0.3.0**. Newer HarnessTrim builds can
|
|
218
|
+
be accepted without another hard-coded version bump when their executable version matches their
|
|
219
|
+
machine-readable `capabilities` version and the semantic surface/write-set comparison reports no
|
|
220
|
+
drift.
|
|
221
|
+
|
|
222
|
+
Provider **package updates are separate from harness configuration writes**. `token-harness update`
|
|
223
|
+
can replace a reviewed provider target without requiring an exact historical Claude/Codex fixture
|
|
224
|
+
for that package version; exact compatibility rows still gate any later managed agent-config
|
|
225
|
+
mutation. HarnessTrim updates use its reviewed npm channel and capture the previous global version
|
|
226
|
+
for rollback. On native Windows RTK still prefers WinGet, but when that catalog is behind the
|
|
227
|
+
reviewed 0.49.0 target Token Harness can fall back to the exact official GitHub Windows x64 release:
|
|
228
|
+
it verifies GitHub's published SHA-256, replaces only the uniquely resolved `rtk.exe`, verifies the
|
|
229
|
+
new version, and restores and re-verifies the previous bytes on failure. This package-only fallback
|
|
230
|
+
does not widen RFC 0009 or grant permission to mutate agent configuration.
|
|
231
|
+
|
|
232
|
+
RTK has no equivalent machine-readable capability endpoint, so releases newer than the explicitly
|
|
233
|
+
reviewed RTK set remain visible as `unknown-newer` until their consumed contract is checked. See
|
|
234
|
+
[docs/provider-version-compatibility.md](docs/provider-version-compatibility.md).
|
|
235
|
+
|
|
236
|
+
Current experimental candidates include Headroom, mcptoon and GitNexus. Detection or a promising
|
|
237
|
+
benchmark is not enough for promotion. Their campaign assessment is structured evidence for the
|
|
238
|
+
selection gate, not an activation or promotion decision. A candidate must pass structured promotion
|
|
239
|
+
readiness across benchmark capability, category fit, selection evidence, real activation
|
|
240
|
+
verification, managed lifecycle, compatibility/reversibility, project maturity and combined-stack
|
|
241
|
+
validation. Broader context owners also require an explicit admission decision.
|
|
242
|
+
|
|
243
|
+
See [docs/optimizer-priorities.md](docs/optimizer-priorities.md) and
|
|
244
|
+
[RFC 0027](docs/rfcs/0027-optimization-stack-manager.md).
|
|
245
|
+
|
|
246
|
+
## Stable-stack operating model
|
|
247
|
+
|
|
248
|
+
The intended lifecycle is:
|
|
249
|
+
|
|
250
|
+
```text
|
|
251
|
+
discover -> evaluate -> recommend -> install/configure -> verify -> measure
|
|
252
|
+
-> monitor -> update/re-evaluate -> rollback/uninstall
|
|
195
253
|
```
|
|
196
254
|
|
|
197
|
-
|
|
255
|
+
A healthy deterministic component should mostly be left alone. Re-evaluation is useful when an
|
|
256
|
+
agent/optimizer changes version, configuration drift appears, measured value deteriorates, quality
|
|
257
|
+
regresses, workload shape changes materially, or a credible better candidate appears.
|
|
198
258
|
|
|
199
|
-
##
|
|
259
|
+
## Use Token Harness with an AI agent
|
|
200
260
|
|
|
201
|
-
|
|
202
|
-
|
|
261
|
+
The browser remains the primary human interface, but the repository also includes a portable Agent
|
|
262
|
+
Skill at [`skills/token-harness/SKILL.md`](skills/token-harness/SKILL.md).
|
|
203
263
|
|
|
204
|
-
|
|
264
|
+
If you prefer, you can ask Claude Code or Codex to help with installation and inspection. For
|
|
265
|
+
example:
|
|
205
266
|
|
|
206
|
-
|
|
207
|
-
|
|
208
|
-
|
|
209
|
-
|
|
210
|
-
`uninstall --yes` are the explicit CLI configuration-changing forms; the guided UI uses
|
|
211
|
-
a reviewed preview and explicit **Approve and apply** instead;
|
|
212
|
-
- plans are checked again immediately before they are applied;
|
|
213
|
-
- existing files are backed up before a managed write;
|
|
214
|
-
- only exact Token Harness-owned entries are removed by `uninstall`;
|
|
215
|
-
- newer or untested combinations are reported, not guessed;
|
|
216
|
-
- an available provider update outside reviewed compatibility is kept out rather than
|
|
217
|
-
forced, and the installed working version stays in place;
|
|
218
|
-
- the guided app binds only to 127.0.0.1 and protects its fixed local controls with exact
|
|
219
|
-
Host/Origin checks, a per-process anti-forgery token and single-use approval tickets;
|
|
220
|
-
- the legacy read-only dashboard and external status seam remain read-only;
|
|
221
|
-
- source code, prompts, command contents, credentials, and cookies are not sent to a
|
|
222
|
-
Token Harness service.
|
|
223
|
-
|
|
224
|
-
Plans, receipts, metrics, and backups stay in the local Token Harness state directory.
|
|
225
|
-
See [RFC 0013](docs/rfcs/0013-guided-local-experience.md) for the local browser trust boundary,
|
|
226
|
-
[RFC 0004](docs/rfcs/0004-safety-and-installation.md) for the execution model and
|
|
227
|
-
[RFC 0006](docs/rfcs/0006-cli-contract.md) for CLI/JSON guarantees.
|
|
228
|
-
|
|
229
|
-
## Supported optimizations
|
|
230
|
-
|
|
231
|
-
Token Harness can detect and measure several independent local tools:
|
|
232
|
-
|
|
233
|
-
| Provider | Purpose | Management |
|
|
234
|
-
| --- | --- | --- |
|
|
235
|
-
| [RTK](https://github.com/rtk-ai/rtk) | Shell-command rewriting and output reduction | Managed only for reviewed combinations |
|
|
236
|
-
| [HarnessTrim](https://github.com/giuliastro/HarnessTrim) | Deterministic reducers and harness adapters | Managed only for reviewed combinations |
|
|
237
|
-
| [cclimits](https://github.com/cruzanstx/cclimits) | Optional live/local quota companion | Read-only; never installed automatically |
|
|
238
|
-
| [ccusage](https://github.com/ccusage/ccusage) | Local usage history | Read-only; never installed automatically |
|
|
267
|
+
```text
|
|
268
|
+
Install the latest Token Harness, open it, inspect my coding-agent setup, and explain any proposed
|
|
269
|
+
change before applying it. Do not apply configuration changes without my approval.
|
|
270
|
+
```
|
|
239
271
|
|
|
240
|
-
|
|
241
|
-
|
|
272
|
+
The skill is deliberately thin: Token Harness remains the deterministic stack/evidence controller.
|
|
273
|
+
The AI does not bypass preview, compatibility checks or explicit approval.
|
|
242
274
|
|
|
243
|
-
|
|
244
|
-
|
|
245
|
-
|
|
275
|
+
The app can also preview enabling that guidance in supported user-level Agent Skills locations.
|
|
276
|
+
Existing custom skill directories are not silently overwritten or adopted. See
|
|
277
|
+
[RFC 0023](docs/rfcs/0023-guided-agent-skill-install.md).
|
|
246
278
|
|
|
247
|
-
## Advanced
|
|
279
|
+
## Advanced CLI
|
|
248
280
|
|
|
249
|
-
Most people do not need
|
|
281
|
+
Most people do not need these commands. They remain available for automation, debugging and the
|
|
282
|
+
browser controller itself.
|
|
250
283
|
|
|
251
284
|
| Command | Purpose | Changes agent/project config? |
|
|
252
285
|
| --- | --- | --- |
|
|
253
|
-
| `doctor` | Detect
|
|
286
|
+
| `doctor` | Detect agents, providers, versions and problems | No |
|
|
254
287
|
| `budget` | Read authoritative/reported allowance windows | No |
|
|
255
|
-
| `context` | Inspect model settings, instructions
|
|
288
|
+
| `context` | Inspect model settings, instructions and MCP exposure | No |
|
|
256
289
|
| `mcp` | Focus on MCP server/tool health | No |
|
|
257
290
|
| `history` | Summarize local usage through an installed ccusage | No |
|
|
258
291
|
| `plan` | Prepare exact supported changes | No; stores local plan state |
|
|
259
292
|
| `apply` | Apply a reviewed stored plan | Yes, only with `--yes` |
|
|
260
293
|
| `verify` | Check the declared integration tier | No |
|
|
261
294
|
| `metrics` | Report attributable reducer savings | No |
|
|
262
|
-
| `status` | Report pipelines, drift
|
|
263
|
-
| `update` | Check/update
|
|
295
|
+
| `status` | Report pipelines, drift and importer modes | No |
|
|
296
|
+
| `update` | Check/update reviewed provider packages | Yes, only with `--yes` |
|
|
264
297
|
| `rollback` | Restore the latest transaction snapshot | Yes, only with `--yes` |
|
|
265
298
|
| `uninstall` | Remove owned integration entries | Yes, only with `--yes` |
|
|
266
299
|
| `schedule` | Compare Claude Code and Codex using available evidence | No |
|
|
267
|
-
| `handoff` | Build a bounded cross-
|
|
300
|
+
| `handoff` | Build a bounded cross-agent handoff | No |
|
|
268
301
|
| `benchmark*`, `transfer*` | Capture and compare empirical evidence | Local state only |
|
|
269
302
|
|
|
270
|
-
|
|
303
|
+
Need stable machine-readable output? Add `--json`. Need the evidence behind a human summary? Add
|
|
304
|
+
`--verbose`.
|
|
305
|
+
|
|
306
|
+
The older automation contracts remain available. `ui --json` preserves its existing schema-1
|
|
307
|
+
report; `ui --read-only` opens the legacy read-only UI; `ui --no-open` starts the guided app without
|
|
308
|
+
launching a browser.
|
|
271
309
|
|
|
272
|
-
|
|
273
|
-
|
|
310
|
+
### Evaluating an experimental candidate
|
|
311
|
+
|
|
312
|
+
For normal use, open **Setup -> Experimental tools** and choose **Run standard evaluation**. The app
|
|
313
|
+
keeps a resumable campaign ID for each candidate/harness pair and reads campaign progress, assessment
|
|
314
|
+
and the exact **Next** step directly in the browser. Use **Start baseline capture** or **Start optimized
|
|
315
|
+
capture**, run the requested task in the selected coding agent, then choose **Record outcome** and
|
|
316
|
+
enter the quality/attempt values you actually observed. Normal use no longer requires copying
|
|
317
|
+
`benchmark-start` or `benchmark-finish` commands into a terminal.
|
|
318
|
+
|
|
319
|
+
The equivalent advanced CLI flow starts by asking the campaign engine for its current state:
|
|
274
320
|
|
|
275
321
|
```sh
|
|
276
|
-
token-harness
|
|
277
|
-
|
|
322
|
+
token-harness benchmark-matrix \
|
|
323
|
+
--benchmark-id gitnexus-codex-eval-1 \
|
|
324
|
+
--candidate gitnexus \
|
|
325
|
+
--harness codex
|
|
278
326
|
```
|
|
279
327
|
|
|
280
|
-
|
|
281
|
-
|
|
282
|
-
|
|
283
|
-
|
|
284
|
-
|
|
328
|
+
Follow only the **Next** command printed by that report, complete the task honestly, then rerun the
|
|
329
|
+
same `benchmark-matrix` command. Before an optimized run, enable the candidate through its own
|
|
330
|
+
documented workflow. Token Harness records the experiment target but does not treat attribution—or
|
|
331
|
+
the browser acknowledgement—as proof that the candidate was active. For GitNexus, new benchmark
|
|
332
|
+
receipts can additionally verify activation from the harness-native MCP inventory when the GitNexus
|
|
333
|
+
server is observed usable at both task boundaries; missing or ambiguous runtime evidence stays
|
|
334
|
+
unverified.
|
|
285
335
|
|
|
286
|
-
|
|
287
|
-
|
|
288
|
-
|
|
289
|
-
harness, but only when that candidate has enough conservative accepted-task capacity and passes the
|
|
290
|
-
existing quality, pace, availability, and transfer checks.
|
|
336
|
+
The selection assessment can become decision-ready after enough evidence across task classes, but it
|
|
337
|
+
still cannot promote a candidate by itself. The remaining lifecycle and combined-stack gates must be
|
|
338
|
+
satisfied separately.
|
|
291
339
|
|
|
292
|
-
###
|
|
340
|
+
### Workload-aware allowance planning
|
|
293
341
|
|
|
294
|
-
|
|
342
|
+
If you explicitly know the remaining backlog, the advanced CLI can reason about whether that work
|
|
343
|
+
fits the currently observed allowance:
|
|
295
344
|
|
|
296
345
|
```sh
|
|
297
|
-
token-harness
|
|
298
|
-
|
|
346
|
+
token-harness optimize --harness codex --task standard --tasks-left 5
|
|
347
|
+
token-harness schedule --current codex --candidate claude --task-class standard --tasks-left 5
|
|
299
348
|
```
|
|
300
349
|
|
|
301
|
-
|
|
302
|
-
each proposed task's project-local p75 cost against the same shared five-hour and weekly allowance
|
|
303
|
-
of that harness, then returns `stay`, `split`, `switch`, `shortfall`, or
|
|
304
|
-
`insufficient-evidence`. Candidate assignments require at least three coherent quality-gated
|
|
305
|
-
observations for the exact task class plus complete five-hour and weekly capacity evidence.
|
|
306
|
-
Unproven work remains visibly unallocated.
|
|
350
|
+
For a mixed queued workload:
|
|
307
351
|
|
|
308
|
-
|
|
309
|
-
|
|
310
|
-
|
|
311
|
-
|
|
352
|
+
```sh
|
|
353
|
+
token-harness schedule --current codex --candidate claude \
|
|
354
|
+
--workload mechanical=2,standard=3,hard=1
|
|
355
|
+
```
|
|
312
356
|
|
|
313
|
-
|
|
314
|
-
|
|
315
|
-
|
|
357
|
+
Token Harness does not infer remaining tasks from session length, local tokens or raw provider
|
|
358
|
+
percentages. If the required benchmark/allowance evidence is incomplete, capacity remains unknown.
|
|
359
|
+
See [RFC 0020](docs/rfcs/0020-workload-aware-allowance.md) and
|
|
360
|
+
[RFC 0021](docs/rfcs/0021-mixed-workload-allocation.md).
|
|
316
361
|
|
|
317
|
-
|
|
362
|
+
### Applying native recommendations from the CLI
|
|
318
363
|
|
|
319
|
-
`optimize` remains read-only.
|
|
364
|
+
`optimize` remains read-only. The explicit CLI path is review then apply:
|
|
320
365
|
|
|
321
366
|
```sh
|
|
322
367
|
token-harness plan --harness claude --native-policy --task mechanical --profile economy
|
|
323
368
|
token-harness apply --plan <printed-plan-id> --yes
|
|
324
369
|
```
|
|
325
370
|
|
|
326
|
-
|
|
327
|
-
should not have to repeat `--harness`, `--provider`, `--native-policy`, `--task` or `--profile`.
|
|
328
|
-
Run it from the same project as `plan`. Conflicting explicit selectors are rejected, and
|
|
329
|
-
actual version, ownership and configuration changes still invalidate the plan. Existing
|
|
330
|
-
schema-1 plans remain usable; only their approved actions can execute.
|
|
371
|
+
For normal use, prefer the browser workflow.
|
|
331
372
|
|
|
332
|
-
|
|
333
|
-
Claude Code 2.1.261 build. It does not change model, authentication, hooks, endpoint or billing.
|
|
334
|
-
`max` is never persisted. Project/local/ancestor settings, custom configuration roots and
|
|
335
|
-
known environment/thinking overrides block the change rather than being overwritten. The
|
|
336
|
-
preference affects future sessions unless overridden: reopen Claude and check `/effort`.
|
|
337
|
-
This is not evidence of a running session's effective effort or a guaranteed quota saving.
|
|
373
|
+
## Safety and privacy
|
|
338
374
|
|
|
339
|
-
|
|
340
|
-
verbosity fields through native `config/batchWrite`; project/profile overrides remain yours.
|
|
341
|
-
`rollback --yes` restores the complete pre-change files. `uninstall --yes` removes only owned
|
|
342
|
-
changes and restores a prior Claude effort preference without undoing unrelated later edits.
|
|
375
|
+
Token Harness is conservative by design:
|
|
343
376
|
|
|
344
|
-
|
|
377
|
+
- opening the app and normal read-only commands do not change agent/project configuration;
|
|
378
|
+
- a browser configuration mutation requires preview and explicit approval;
|
|
379
|
+
- guided candidate capture buttons write only bounded local benchmark state, are CSRF-protected and
|
|
380
|
+
must still match the campaign engine's current step immediately before the write;
|
|
381
|
+
- a browser activation acknowledgement is never treated as verified candidate activation;
|
|
382
|
+
- CLI mutations require their explicit `--yes` form;
|
|
383
|
+
- plans are checked again immediately before apply;
|
|
384
|
+
- existing files are backed up before a managed write;
|
|
385
|
+
- only exact Token Harness-owned entries are removed by uninstall;
|
|
386
|
+
- newer or untested combinations are reported rather than guessed;
|
|
387
|
+
- an available provider update outside reviewed package compatibility is kept out rather than forced;
|
|
388
|
+
- provider package replacement does not bypass the stricter compatibility gate for harness config writes;
|
|
389
|
+
- the guided app binds only to `127.0.0.1` and protects local controls with Host/Origin checks, a
|
|
390
|
+
per-process anti-forgery token and single-use approval tickets;
|
|
391
|
+
- source code, prompts, command contents, credentials and cookies are not sent to a Token Harness
|
|
392
|
+
service.
|
|
345
393
|
|
|
346
|
-
|
|
394
|
+
Plans, receipts, metrics and backups stay in the local Token Harness state directory.
|
|
395
|
+
|
|
396
|
+
See [RFC 0013](docs/rfcs/0013-guided-local-experience.md),
|
|
397
|
+
[RFC 0004](docs/rfcs/0004-safety-and-installation.md), and
|
|
398
|
+
[RFC 0006](docs/rfcs/0006-cli-contract.md).
|
|
347
399
|
|
|
348
|
-
|
|
349
|
-
flags, cannot find Python, has no usable Claude session, reports an expired session, or returns
|
|
350
|
-
an unsupported source. It does not expose credentials, raw companion errors or private paths.
|
|
400
|
+
## Updating, checking and undoing
|
|
351
401
|
|
|
352
|
-
|
|
353
|
-
zero-configuration support and the read-only flags. The latest GitHub Release listing is older
|
|
354
|
-
and is not evidence of what npm ships. To check the same path Token Harness uses:
|
|
402
|
+
Update Token Harness itself:
|
|
355
403
|
|
|
356
404
|
```sh
|
|
357
|
-
npm
|
|
358
|
-
cclimits --claude --json --no-cache-write --no-stale-fallback
|
|
359
|
-
token-harness budget --harness claude --verbose
|
|
405
|
+
npm install --global token-harness@latest
|
|
360
406
|
```
|
|
361
407
|
|
|
362
|
-
|
|
363
|
-
|
|
364
|
-
A fresh local Claude cache is shown as **cached**, never promoted to live quota pacing.
|
|
365
|
-
A missing observation is not zero remaining allowance. Never paste credentials to debug it.
|
|
366
|
-
|
|
367
|
-
### Codex is configured but its hook does not run
|
|
408
|
+
Provider checks and reviewed updates are in **Setup -> Checks and maintenance**. From the advanced
|
|
409
|
+
CLI, first preview and then explicitly apply the provider-package updates:
|
|
368
410
|
|
|
369
|
-
|
|
370
|
-
|
|
371
|
-
|
|
372
|
-
|
|
373
|
-
`config-only` until attributable runtime evidence exists.
|
|
411
|
+
```sh
|
|
412
|
+
token-harness update
|
|
413
|
+
token-harness update --yes
|
|
414
|
+
```
|
|
374
415
|
|
|
375
|
-
|
|
416
|
+
`update` replaces only installed providers whose target is inside the reviewed provider-package
|
|
417
|
+
policy. HarnessTrim uses npm and captures the previous global version for rollback. On native
|
|
418
|
+
Windows RTK prefers WinGet; when WinGet cannot yet reach the reviewed target, Token Harness can use
|
|
419
|
+
the verified official GitHub Windows x64 release fallback described above. That fallback verifies the
|
|
420
|
+
published digest and post-update version and restores the previous executable on failure.
|
|
376
421
|
|
|
377
|
-
|
|
422
|
+
Remove only Token Harness-owned integration entries:
|
|
378
423
|
|
|
379
424
|
```sh
|
|
380
|
-
|
|
381
|
-
npm list --global token-harness
|
|
425
|
+
token-harness uninstall --yes
|
|
382
426
|
```
|
|
383
427
|
|
|
384
|
-
|
|
385
|
-
|
|
386
|
-
### Setup needs attention
|
|
387
|
-
|
|
388
|
-
Run the single command it prints. For technical evidence:
|
|
428
|
+
Restore complete files from the latest committed transaction snapshot:
|
|
389
429
|
|
|
390
430
|
```sh
|
|
391
|
-
token-harness
|
|
431
|
+
token-harness rollback --yes
|
|
392
432
|
```
|
|
393
433
|
|
|
394
|
-
|
|
395
|
-
you
|
|
434
|
+
`rollback` is whole-file time travel and can also revert later manual edits to those files. Prefer
|
|
435
|
+
`uninstall` when you only want to remove Token Harness-owned entries.
|
|
396
436
|
|
|
397
|
-
|
|
437
|
+
## Troubleshooting
|
|
398
438
|
|
|
399
|
-
|
|
400
|
-
newer provider release but does not yet have reviewed compatibility evidence for the
|
|
401
|
-
active provider × harness × platform combination. Keep using the installed version; no
|
|
402
|
-
manual upgrade is required.
|
|
439
|
+
### Claude allowance is unavailable
|
|
403
440
|
|
|
404
|
-
|
|
441
|
+
The app explains whether the optional cclimits companion is missing, too old for safe read-only
|
|
442
|
+
flags, cannot find Python, has no usable Claude session, reports an expired session, or returns an
|
|
443
|
+
unsupported source. It does not expose credentials, raw companion errors or private paths.
|
|
405
444
|
|
|
406
|
-
|
|
445
|
+
For the same technical evidence in the terminal:
|
|
407
446
|
|
|
408
447
|
```sh
|
|
409
|
-
|
|
448
|
+
npm list --global cclimits
|
|
449
|
+
cclimits --claude --json --no-cache-write --no-stale-fallback
|
|
450
|
+
token-harness budget --harness claude --verbose
|
|
410
451
|
```
|
|
411
452
|
|
|
412
|
-
|
|
413
|
-
the two states separately.
|
|
414
|
-
|
|
415
|
-
## Updating or undoing
|
|
453
|
+
A missing observation is not zero remaining allowance. Never paste credentials to debug it.
|
|
416
454
|
|
|
417
|
-
|
|
455
|
+
### `token-harness` is not found
|
|
418
456
|
|
|
419
457
|
```sh
|
|
420
|
-
|
|
421
|
-
token-harness
|
|
458
|
+
node --version
|
|
459
|
+
npm list --global token-harness
|
|
422
460
|
```
|
|
423
461
|
|
|
424
|
-
|
|
462
|
+
Node must be at least 22.13. Reopen the terminal after installation if needed.
|
|
463
|
+
|
|
464
|
+
### Setup or verification needs attention
|
|
465
|
+
|
|
466
|
+
Use the action shown in Dashboard or Setup. For technical evidence:
|
|
425
467
|
|
|
426
468
|
```sh
|
|
427
|
-
token-harness
|
|
469
|
+
token-harness doctor --verbose
|
|
470
|
+
token-harness verify --verbose
|
|
428
471
|
```
|
|
429
472
|
|
|
430
|
-
|
|
473
|
+
Do not force an unsupported plan. A newer version outside reviewed compatibility is normally a
|
|
474
|
+
safety limitation, not a reason to overwrite the known-working installation.
|
|
475
|
+
|
|
476
|
+
## Run the current source
|
|
477
|
+
|
|
478
|
+
From an existing clone:
|
|
431
479
|
|
|
432
480
|
```sh
|
|
433
|
-
|
|
481
|
+
npm start
|
|
434
482
|
```
|
|
435
483
|
|
|
436
|
-
|
|
437
|
-
files. Prefer `uninstall` when you only want to remove Token Harness-owned entries.
|
|
438
|
-
|
|
439
|
-
## Develop from source
|
|
484
|
+
For a fresh clone:
|
|
440
485
|
|
|
441
486
|
```sh
|
|
442
487
|
git clone https://github.com/giuliastro/token-harness.git
|
|
443
488
|
cd token-harness
|
|
444
|
-
|
|
445
|
-
|
|
446
|
-
pnpm typecheck
|
|
447
|
-
pnpm lint
|
|
448
|
-
pnpm test
|
|
449
|
-
pnpm build
|
|
450
|
-
pnpm smoke
|
|
451
|
-
pnpm package
|
|
452
|
-
pnpm smoke:install
|
|
489
|
+
npx --yes pnpm@10.33.4 install --frozen-lockfile
|
|
490
|
+
npm start
|
|
453
491
|
```
|
|
454
492
|
|
|
455
|
-
|
|
456
|
-
|
|
493
|
+
This runs the clone. An unmerged branch or unpublished `main` change is not automatically available
|
|
494
|
+
through `token-harness@latest`.
|
|
457
495
|
|
|
458
|
-
##
|
|
496
|
+
## Development
|
|
459
497
|
|
|
460
|
-
|
|
461
|
-
|
|
498
|
+
```sh
|
|
499
|
+
git clone https://github.com/giuliastro/token-harness.git
|
|
500
|
+
cd token-harness
|
|
501
|
+
npx --yes pnpm@10.33.4 install --frozen-lockfile
|
|
502
|
+
npx --yes pnpm@10.33.4 typecheck
|
|
503
|
+
npx --yes pnpm@10.33.4 lint
|
|
504
|
+
npx --yes pnpm@10.33.4 test
|
|
505
|
+
npx --yes pnpm@10.33.4 build
|
|
506
|
+
npx --yes pnpm@10.33.4 smoke
|
|
507
|
+
npx --yes pnpm@10.33.4 package
|
|
508
|
+
npx --yes pnpm@10.33.4 smoke:install
|
|
509
|
+
```
|
|
462
510
|
|
|
463
|
-
|
|
511
|
+
Using `corepack enable` is optional. On a system-wide Windows Node installation it can require
|
|
512
|
+
administrator permission to modify `C:\Program Files\nodejs`; the `npx pnpm@10.33.4` form above
|
|
513
|
+
does not require that Corepack shim write.
|
|
464
514
|
|
|
465
|
-
|
|
466
|
-
|
|
467
|
-
|
|
468
|
-
|
|
515
|
+
Before changing public behavior or architecture, read
|
|
516
|
+
[RFC 0027](docs/rfcs/0027-optimization-stack-manager.md),
|
|
517
|
+
[docs/optimizer-priorities.md](docs/optimizer-priorities.md),
|
|
518
|
+
[docs/release-readiness.md](docs/release-readiness.md), [PLAN.md](PLAN.md), and the accepted
|
|
519
|
+
[RFCs](docs/rfcs).
|
|
469
520
|
|
|
470
|
-
|
|
471
|
-
of recorded changed outputs immediately beside it. An estimate says **"Estimated"**. This
|
|
472
|
-
percentage describes only those recorded outputs, not your whole coding session, subscription
|
|
473
|
-
allowance or money. Provider rows remain separate; negative results and errors remain visible.
|
|
521
|
+
## License
|
|
474
522
|
|
|
475
|
-
|
|
476
|
-
|
|
477
|
-
text post and choose a community yourself. **Copy for Discord** prepares a message to paste
|
|
478
|
-
in your chosen channel. **Save image** creates a PNG you can attach yourself. Nothing is
|
|
479
|
-
posted or uploaded automatically, and sharing excludes private paths, code, prompts and
|
|
480
|
-
account/allowance information. An open share preview stays fixed even if readings update.
|
|
523
|
+
[Apache License 2.0](LICENSE). Referenced provider tools are independent projects with their own
|
|
524
|
+
licenses.
|