token-harness 0.1.11 → 0.1.13

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (4) hide show
  1. package/README.md +79 -122
  2. package/package.json +1 -1
  3. package/sbom.json +3 -3
  4. package/token-harness.mjs +4715 -1333
package/README.md CHANGED
@@ -3,8 +3,8 @@
3
3
  **Build, verify and measure an optimization stack for Claude Code and Codex.**
4
4
 
5
5
  Token Harness is a local **optimization stack manager**. It checks your coding agents, manages the
6
- optimization components it can safely own, keeps experimental candidates separate, verifies the
7
- result, and reports savings only when it has evidence to support them.
6
+ optimization components it can safely own, keeps evaluation evidence separate from managed lifecycle,
7
+ verifies the result, and reports savings only when it has evidence to support them.
8
8
 
9
9
  It is not another coding agent and it does not replace specialized projects such as RTK or
10
10
  HarnessTrim.
@@ -29,117 +29,67 @@ of CLI commands to memorize.
29
29
 
30
30
  ### First run
31
31
 
32
- 1. Open **Dashboard** and let Token Harness inspect the current setup.
33
- 2. If setup is incomplete, choose **Open setup**.
34
- 3. In **Setup**, work from top to bottom:
35
- - Coding agents
36
- - Managed optimizers
37
- - Experimental tools
38
- - Optional agent tuning
39
- - Checks and maintenance
40
- 4. For Claude Code or Codex, choose **Review setup** when a managed setup is available.
41
- 5. Read the exact proposed changes. **Apply reviewed setup** appears only when there is a concrete
42
- safe plan to apply.
43
- 6. Keep using Claude Code or Codex normally.
44
- 7. Open **Results** when you want to see what Token Harness can actually prove.
45
-
46
- Opening the app does not change your configuration. Read-only checks stay read-only, and a managed
47
- write requires an explicit review and approval.
48
-
49
- ## The three views
50
-
51
- ### Dashboard
52
-
53
- Dashboard answers the questions that matter first:
54
-
55
- - is my setup ready;
56
- - which coding agents and managed optimizers are active;
57
- - what should I do next;
58
- - what value has actually been measured;
59
- - whether quality or an integration needs attention.
60
-
61
- The headline cards deliberately distinguish measured evidence from unknown values. Missing evidence
62
- is never shown as zero savings.
63
-
64
- ### Setup
65
-
66
- Setup is one ordered workflow instead of a collection of unrelated actions.
67
-
68
- **1. Coding agents**
69
-
70
- Token Harness currently supports guided setup for Claude Code and Codex. Agent details also show
71
- useful read-only allowance and connected-tool observations when available.
72
-
73
- **2. Managed optimizers**
74
-
75
- The managed stack currently consists of **RTK + HarnessTrim** on individually reviewed
76
- combinations. Token Harness tracks the exact combined provider set separately: if no combined-stack
77
- review is recorded, Setup says so and keeps the stack incomplete rather than inferring compatibility
78
- from healthy individual checks. Token Harness can prepare their integration transactionally, show the
79
- exact plan, apply it only after approval, verify it, and remove only configuration it owns.
80
-
81
- For maintainers validating the combined stack, `token-harness stack-review --json` captures the exact
82
- configured provider versions and managed harness sets and reuses the existing passive `verify`
83
- evidence for exact provider/harness pairs. Runtime evidence is credited only when it can be
84
- attributed to that harness: HarnessTrim uses its native event harness field, while provider-wide
85
- telemetry can be attributed by exclusion only when one harness is wired. With multiple harnesses, an
86
- unattributable receipt stays **Unavailable** instead of being copied across rows. A `not-exercised`
87
- result does not promise that ordinary agent use will create a receipt: the provider must record a
88
- qualifying operation, and for reducers that means a real reduction. `stack-review` does not run an
89
- active canary or spend a model call, and it never makes the compatibility decision itself. The
90
- shipped combined-review registry stays empty until the captured configuration has enough real
91
- verification evidence, has been benchmarked together, and has been deliberately reviewed; see
92
- `docs/combined-stack-reviews.md`.
93
-
94
- **3. Experimental tools**
95
-
96
- Headroom, mcptoon and GitNexus are visible as candidates, not silently promoted dependencies. Their
97
- cards distinguish CLI installation, candidate-side activation/evaluation, Token Harness benchmark
98
- evidence and promotion-readiness gates. External install or activation commands are shown for
99
- review; Token Harness does not silently execute package managers, activate wrappers, index
100
- repositories or register MCP servers.
101
-
102
- Choose **Run standard evaluation** to start or resume a paired candidate campaign. Campaign state is
103
- scoped to both the candidate and the selected harness, so Claude Code and Codex evidence cannot be
104
- mixed accidentally. The browser reads the campaign directly and shows **Progress**, the current
105
- selection signal, whether the evidence is **Decision ready**, the number of evidence-bearing pairs
106
- and the exact **Next** step.
107
-
108
- For normal use, the browser can now start and finish the local benchmark capture itself. You still
109
- run the actual task in Claude Code or Codex. When the task finishes, record the quality result,
110
- attempt count and failed-attempt count you actually observed. Before an optimized capture, Token
111
- Harness requires you to acknowledge that you enabled the candidate through its own documented
112
- workflow. That acknowledgement is **not activation verification** and is never treated as promotion
113
- evidence. Every browser capture action is matched against the campaign engine's current step before
114
- it can write local benchmark state, so a stale tab cannot advance a different step.
115
-
116
- `benchmark-matrix`, `benchmark-start` and `benchmark-finish` remain available as advanced terminal
117
- fallbacks for debugging or automation. The browser does not run the coding task, install or activate
118
- a candidate, or claim that candidate attribution proves activation.
119
-
120
- A campaign selection assessment can report `insufficient-evidence`, `promising`, `mixed` or
121
- `negative`, plus whether the evidence is decision-ready. **Decision-ready is not promotion-ready.**
122
- The campaign shows the same promotion-review gate count and next gate as the candidate card, using
123
- the current local observation rather than starting another environment scan. Activation
124
- verification, managed lifecycle, compatibility/reversibility, project maturity and combined-stack
125
- validation remain separate gates.
32
+ The first screen is **Overview**. There is no separate Setup page to learn.
33
+
34
+ 1. Token Harness detects Claude Code and Codex.
35
+ 2. Each detected coding agent says either **Ready** or **Setup incomplete**.
36
+ 3. If setup is incomplete, choose **Finish setup** on that agent. Token Harness prepares the
37
+ recommended **RTK + HarnessTrim** baseline and shows the exact safe plan before anything changes.
38
+ 4. The **Optimizers** section keeps the recommended baseline separate from optional tools
39
+ (mcptoon, GitNexus and Headroom). When a tool needs an action, that action appears on the same
40
+ card as the status it refers to.
41
+ 5. **Health and updates** is maintenance, not another onboarding checklist. Normal setup performs its
42
+ own safety checks. Use **Re-check health** for troubleshooting and **Check for updates** when you
43
+ want to inspect provider versions. If a reviewed update is available, the same dialog offers
44
+ **Install updates** after showing the versions.
45
+ 6. Keep using Claude Code or Codex normally. Open **Results** when you want detailed evidence.
46
+
47
+ Opening the app does not change your configuration. A configuration or software change is always
48
+ previewed first and requires an explicit review and approval.
49
+
50
+ ## The two views
51
+
52
+ ### Overview
53
+
54
+ Overview is both the first-run screen and the normal status screen. It keeps the two product entities
55
+ separate:
56
+
57
+ - **Coding agents** — currently Claude Code and Codex. Detection only means Token Harness can see the
58
+ agent; first-run setup is complete for that agent only when both RTK and HarnessTrim are connected.
59
+ - **Optimizers** — RTK and HarnessTrim are the recommended baseline. mcptoon, GitNexus and Headroom
60
+ are optional reviewed integrations with narrower prerequisites and compatibility boundaries.
61
+
62
+ The status at the top always answers what to do next. States such as **Setup incomplete**,
63
+ **Installed · not connected**, **Needs attention**, or **Update available** have their action beside
64
+ the affected agent or optimizer instead of in a separate action list.
65
+
66
+ Overview also contains a compact **Measured impact** summary. Missing evidence is shown as unknown,
67
+ never as zero savings.
68
+
69
+ Advanced agent details and reasoning preferences are collapsed because they are not required for
70
+ first-run optimizer setup.
126
71
 
127
- Choose **Compare evaluation evidence** when you want one read-only view of campaigns you already
128
- started. The comparison is loaded only on request, never creates a campaign, keeps candidate and
129
- harness ordering fixed, and reports progress, selection signal, decision readiness, runtime
130
- activation evidence and promotion gates without producing a composite score or automatic winner.
131
- Activation is shown separately as **Verified**, **Blocked** or **Unreviewed**, with
132
- verified/blocked/unknown pair counts when campaign evidence exists.
72
+ ### Optimizer lifecycle and evidence
133
73
 
134
- **4. Optional agent tuning**
74
+ The recommended baseline remains **RTK + HarnessTrim** on individually reviewed combinations.
75
+ Token Harness also exposes **mcptoon**, **GitNexus** and **Headroom** as optional managed integrations
76
+ on exact reviewed lifecycle rows. Enabling an optional optimizer does not make it part of the
77
+ production baseline and does not create a savings claim.
135
78
 
136
- Reasoning preferences are separate from optimizer installation. They are persistent agent settings,
137
- not hidden per-task switches, and are changed only through the normal preview/apply flow.
79
+ Token Harness can prepare supported integration changes transactionally, show the exact plan, apply
80
+ it only after approval, verify what the declared tier can verify, and remove only configuration it
81
+ owns. GitNexus is never auto-indexed and its noncommercial license boundary stays visible. Headroom
82
+ remains config-only managed on its reviewed row; Token Harness does not bootstrap uv/Python or start
83
+ wrapper/proxy/deploy flows.
138
84
 
139
- **5. Checks and maintenance**
85
+ For maintainers, `token-harness stack-review --json` captures the exact configured provider versions
86
+ and managed harness sets for combined-stack review. Runtime evidence is credited only when it can be
87
+ attributed to the relevant harness; missing attribution remains unavailable rather than being copied
88
+ across rows.
140
89
 
141
- Read-only integration checks, update checks and safe removal/undo controls live here. An update
142
- outside reviewed compatibility is not forced.
90
+ Evaluation campaigns are an advanced maintainer workflow, not a novice setup step. Normal managed
91
+ setup stays in the unified **Optimization Stack** and historical candidate evidence never turns into
92
+ a savings, compatibility or promotion claim by itself.
143
93
 
144
94
  ### Results
145
95
 
@@ -209,6 +159,9 @@ algorithms into this repository.
209
159
  | --- | --- | --- |
210
160
  | [RTK](https://github.com/rtk-ai/rtk) | Shell/tool output reduction | Managed on reviewed combinations |
211
161
  | [HarnessTrim](https://github.com/giuliastro/HarnessTrim) | Deterministic output/context reduction | Managed first-party integration |
162
+ | mcptoon | MCP discovery / compact manifest guidance | Optional managed integration on exact reviewed 0.7.10 rows; no savings assumed |
163
+ | GitNexus | Repository graph / MCP context | Optional managed Claude integration for already-installed 1.6.12; license review required |
164
+ | Headroom | Local MCP context compression/retrieval | Optional config-only managed Claude/Codex integration for already-installed 0.37.0; package prerequisite stays user-owned |
212
165
  | [cclimits](https://github.com/cruzanstx/cclimits) | Optional Claude allowance evidence | Read-only evidence; not an optimizer |
213
166
  | [ccusage](https://github.com/ccusage/ccusage) | Local usage history | Read-only evidence; never subscription quota |
214
167
 
@@ -233,8 +186,8 @@ RTK has no equivalent machine-readable capability endpoint, so releases newer th
233
186
  reviewed RTK set remain visible as `unknown-newer` until their consumed contract is checked. See
234
187
  [docs/provider-version-compatibility.md](docs/provider-version-compatibility.md).
235
188
 
236
- Current experimental candidates include Headroom, mcptoon and GitNexus. Detection or a promising
237
- benchmark is not enough for promotion. Their campaign assessment is structured evidence for the
189
+ Historical evaluation evidence remains available for mcptoon, GitNexus and Headroom. Detection or a promising
190
+ benchmark is not enough for a production-stack promotion or savings claim. Their campaign assessment is structured evidence for the
238
191
  selection gate, not an activation or promotion decision. A candidate must pass structured promotion
239
192
  readiness across benchmark capability, category fit, selection evidence, real activation
240
193
  verification, managed lifecycle, compatibility/reversibility, project maturity and combined-stack
@@ -307,31 +260,34 @@ The older automation contracts remain available. `ui --json` preserves its exist
307
260
  report; `ui --read-only` opens the legacy read-only UI; `ui --no-open` starts the guided app without
308
261
  launching a browser.
309
262
 
310
- ### Evaluating an experimental candidate
263
+ ### Evaluation evidence (advanced / maintainers)
311
264
 
312
- For normal use, open **Setup -> Experimental tools** and choose **Run standard evaluation**. The app
265
+ Evaluation campaigns are an advanced maintainer workflow; managed setup stays in the unified **Optimization Stack**. The app
313
266
  keeps a resumable campaign ID for each candidate/harness pair and reads campaign progress, assessment
314
267
  and the exact **Next** step directly in the browser. Use **Start baseline capture** or **Start optimized
315
268
  capture**, run the requested task in the selected coding agent, then choose **Record outcome** and
316
269
  enter the quality/attempt values you actually observed. Normal use no longer requires copying
317
270
  `benchmark-start` or `benchmark-finish` commands into a terminal.
318
271
 
319
- The equivalent advanced CLI flow starts by asking the campaign engine for its current state:
272
+ The equivalent advanced CLI flow starts by asking the campaign engine for its current state. This
273
+ GitNexus example intentionally uses the only currently reviewed campaign row:
320
274
 
321
275
  ```sh
322
276
  token-harness benchmark-matrix \
323
- --benchmark-id gitnexus-codex-eval-1 \
277
+ --benchmark-id gitnexus-claude-eval-1 \
324
278
  --candidate gitnexus \
325
- --harness codex
279
+ --harness claude
326
280
  ```
327
281
 
328
282
  Follow only the **Next** command printed by that report, complete the task honestly, then rerun the
329
283
  same `benchmark-matrix` command. Before an optimized run, enable the candidate through its own
330
284
  documented workflow. Token Harness records the experiment target but does not treat attribution—or
331
- the browser acknowledgement—as proof that the candidate was active. For GitNexus, new benchmark
332
- receipts can additionally verify activation from the harness-native MCP inventory when the GitNexus
333
- server is observed usable at both task boundaries; missing or ambiguous runtime evidence stays
334
- unverified.
285
+ the browser acknowledgement—as proof that the candidate was active. For GitNexus on the reviewed
286
+ Claude Code `2.1.269` × GitNexus `1.6.12` × native-Linux row, the harness-native MCP inventory can
287
+ prove only that the GitNexus server was available at both task boundaries. That is not proof Claude
288
+ actually called a GitNexus tool, so the activation-verification promotion gate remains blocked
289
+ without a separate reviewed usage witness. See
290
+ [`docs/candidates/gitnexus-real-campaign.md`](docs/candidates/gitnexus-real-campaign.md).
335
291
 
336
292
  The selection assessment can become decision-ready after enough evidence across task classes, but it
337
293
  still cannot promote a candidate by itself. The remaining lifecycle and combined-stack gates must be
@@ -405,8 +361,9 @@ Update Token Harness itself:
405
361
  npm install --global token-harness@latest
406
362
  ```
407
363
 
408
- Provider checks and reviewed updates are in **Setup -> Checks and maintenance**. From the advanced
409
- CLI, first preview and then explicitly apply the provider-package updates:
364
+ Provider checks and reviewed updates are in **Overview -> Health and updates**. The browser first
365
+ checks reviewed channels and, when an installable update exists, offers **Install updates** in the
366
+ same dialog. From the advanced CLI, the preview prints the exact confirmation command:
410
367
 
411
368
  ```sh
412
369
  token-harness update
@@ -463,7 +420,7 @@ Node must be at least 22.13. Reopen the terminal after installation if needed.
463
420
 
464
421
  ### Setup or verification needs attention
465
422
 
466
- Use the action shown in Dashboard or Setup. For technical evidence:
423
+ Use the action shown beside the affected agent or optimizer in Overview. For technical evidence:
467
424
 
468
425
  ```sh
469
426
  token-harness doctor --verbose
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "token-harness",
3
- "version": "0.1.11",
3
+ "version": "0.1.13",
4
4
  "description": "Quota-aware efficiency layer for Claude Code and Codex subscription limits.",
5
5
  "license": "Apache-2.0",
6
6
  "type": "module",
package/sbom.json CHANGED
@@ -1,14 +1,14 @@
1
1
  {
2
2
  "bomFormat": "CycloneDX",
3
3
  "specVersion": "1.5",
4
- "serialNumber": "urn:uuid:01436fc5-44c9-56ef-29b7-e0847e19fad9",
4
+ "serialNumber": "urn:uuid:f53f4600-4e02-4fbe-462d-286326adfd0b",
5
5
  "version": 1,
6
6
  "metadata": {
7
7
  "component": {
8
8
  "type": "application",
9
9
  "bom-ref": "token-harness",
10
10
  "name": "token-harness",
11
- "version": "0.1.11",
11
+ "version": "0.1.13",
12
12
  "description": "Quota-aware efficiency layer for Claude Code and Codex subscription limits.",
13
13
  "licenses": [
14
14
  {
@@ -20,7 +20,7 @@
20
20
  "hashes": [
21
21
  {
22
22
  "alg": "SHA-256",
23
- "content": "01436fc544c956ef29b7e0847e19fad9e439c4cdfccc2517e007156a2b9fb2e0"
23
+ "content": "f53f46004e024fbe462d286326adfd0be1a2e5d5c2678346e1e2e7e39bce853b"
24
24
  }
25
25
  ]
26
26
  },