token-harness 0.1.8 → 0.1.10

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (4) hide show
  1. package/README.md +242 -103
  2. package/package.json +1 -1
  3. package/sbom.json +3 -3
  4. package/token-harness.mjs +8064 -2603
package/README.md CHANGED
@@ -6,119 +6,148 @@ Token Harness checks your coding agents, shows subscription allowance when it ca
6
6
  observed reliably, reduces avoidable context overhead, and recommends useful actions.
7
7
  It runs locally and never presents local token estimates as subscription quota.
8
8
 
9
- ## The easy path
9
+ ## Open it. Approve setup. Keep coding.
10
10
 
11
- You need [Node.js 22.13 or newer](https://nodejs.org/) and at least one signed-in coding
12
- agent such as Claude Code or Codex.
13
-
14
- Open a terminal and paste:
11
+ You need [Node.js 22.13 or newer](https://nodejs.org/) and an installed, signed-in
12
+ Claude Code or Codex. Install Token Harness, then open it:
15
13
 
16
14
  ```sh
17
15
  npm install --global token-harness@latest
18
- token-harness setup
16
+ token-harness
19
17
  ```
20
18
 
21
- That is the whole first-time check. `setup` tells you:
19
+ The browser is now the primary interface. **Set up automatically** checks both agents and
20
+ prepares the supported integration changes. It describes each change in plain language.
21
+ Choose **Approve and apply** to apply the reviewed configuration with backups and verification.
22
+ There are no plan IDs to copy and no daily command sequence to remember.
22
23
 
23
- - which coding agents it found;
24
- - which optimizers are already active;
25
- - whether the current setup works;
26
- - what it changed, if anything;
27
- - exactly one next step.
24
+ Already configured? The app shows your existing integrations without replacing them.
25
+ An absent provider or an unreviewed version combination is explained rather than installed
26
+ or forced silently. Automatic setup covers the reviewed integration paths, not every possible
27
+ provider/version. Token Harness does not install Claude Code or Codex or log you in.
28
28
 
29
- The first run does not change Claude Code, Codex, or project configuration. If Token
30
- Harness finds a supported improvement, it shows the safe plan and may suggest:
31
-
32
- ```sh
33
- token-harness setup --yes
34
- ```
29
+ ### Daily use
35
30
 
36
- `--yes` is always explicit. The change is backed up, applied transactionally, and
37
- verified. Unsupported combinations are left untouched.
31
+ Continue launching `claude` or `codex` as usual. Supported output integrations operate in the
32
+ agent, not in the dashboard. You can close the page and its terminal without disabling those
33
+ integrations. Open `token-harness` whenever you want to see results; the visible dashboard
34
+ imports available provider records and refreshes its readings automatically.
38
35
 
39
- ## Open the dashboard
36
+ **Recorded savings** shows retained history across locally recorded projects, with date bounds,
37
+ provider, measurement class, units, changed-output counts, and before/after values. It does not
38
+ add incompatible provider figures together. Negative results remain visible. No telemetry is
39
+ shown as **not measured**, never a reassuring zero or an invented subscription saving.
40
+ Some provider records may predate Token Harness; locally stored records are not guaranteed
41
+ complete lifetime history. RTK history is imported directly. HarnessTrim project-local records
42
+ must have been imported from their project, or exposed through a configured known metrics path;
43
+ the app does not crawl your disk looking for private projects.
40
44
 
41
- After setup, run:
45
+ For a terminal-only summary, the one command is:
42
46
 
43
47
  ```sh
44
- token-harness ui
48
+ token-harness savings
45
49
  ```
46
50
 
47
- The dashboard opens in your browser and answers three questions in this order:
51
+ Optional windows are `--since 7d` and `--since 30d`. The advanced `metrics` command remains
52
+ project-scoped; opening the app from its installation folder does not change the savings scope.
48
53
 
49
- 1. **Can I work normally right now?**
50
- 2. **What is actually active and useful?**
51
- 3. **Is there one action worth taking?**
54
+ ### Find what you need
52
55
 
53
- It shows active/relevant coding agents first, their optimization providers, observable
54
- allowance windows, and task guidance. Tools that are absent do not get large cards;
55
- secondary detected tools are kept out of the main path.
56
+ The app has three tabs: **Overview** for agents and recorded savings, **Rules & settings**
57
+ for one agent's rules at a time, and **Activity** for checks and guarded undo. Theme follows
58
+ your system; the header also offers light and dark modes.
56
59
 
57
- The page is served only on `127.0.0.1`: no Electron app, account, or cloud service.
60
+ Each rule shows what was actually **observed**, separately from how it works. Actions sit
61
+ next to the relevant state: **Adjust reasoning** opens the supported review/apply flow;
62
+ **Change in Claude/Codex** explains native steps when an automatic write is not available.
63
+ Missing allowance data and measurement records have their own setup/help actions.
58
64
 
59
- ### What happens after the dashboard?
65
+ A saved Claude effort can be displayed even on an unreviewed CLI version. That does not
66
+ admit automatic writes on that version. **No saved preference** means the user field is
67
+ absent, not that reasoning is disabled. A failed read is a separate state with a cause.
68
+ The displayed preference is not a live reading of an already-running session.
60
69
 
61
- Usually: **nothing else. You are done.**
70
+ ### The rules are visible
62
71
 
63
- Token Harness is not a launcher and it does not need to stay between you and your coding
64
- agent. Continue exactly as you normally would:
72
+ **Rules & settings** explains each configured rule: what it does, why it is used,
73
+ its mode, and the evidence available. Automatic integrations, persistent preferences,
74
+ observations and features that are not enabled are explicitly distinguished.
65
75
 
66
- ```sh
67
- claude
68
- codex
69
- opencode
70
- ```
76
+ RTK's supported command integration can reduce output automatically. HarnessTrim can use
77
+ adapters or skills/instructions, depending on the installation; skills-only is not a transparent
78
+ hook, and the agent must actually use the reducer. Configured never means every command was
79
+ intercepted. Provider telemetry and its exact/estimated classification remain separate evidence.
80
+
81
+ **Optional: match reasoning to your work** lets you choose the agent and the type of work
82
+ without learning CLI flags. It previews the actual supported effort/verbosity change, then
83
+ applies only after approval. **This is a persistent preference for future sessions, not an
84
+ automatic per-task switch.** It does not switch models, billing, login, or hook trust. The
85
+ baseline automatic setup never guesses a task or quietly lowers reasoning.
71
86
 
72
- Use whichever of those you already use. RTK, HarnessTrim, or another configured provider
73
- runs automatically through that coding agent's integration. You do **not** need a special
74
- `token-harness run` command.
87
+ **Check integrations** performs the existing integration checks from the UI.
88
+ **Undo last change**, available after an application in that dashboard session, previews a
89
+ whole-file backup restoration. It refuses to undo a newer unrelated transaction. It restores
90
+ only the last successful agent transaction; manual edits to those same files after that
91
+ transaction would also be restored, as the confirmation explains.
75
92
 
76
- You can close the dashboard whenever you want. Use `Ctrl+C` in the terminal to stop its
77
- local web server. Closing it does not disable configured optimizers.
93
+ ### Run the current source
78
94
 
79
- Open it again later with `token-harness ui` when you want a status check. If you do not
80
- want it to open a browser:
95
+ From an existing clone, after installing its dependencies, one command builds and opens the app:
81
96
 
82
97
  ```sh
83
- token-harness ui --no-open
98
+ npm start
84
99
  ```
85
100
 
86
- ## Daily use
101
+ For a fresh clone:
87
102
 
88
- There is no mandatory command loop. These are tools you use when they answer a question:
89
-
90
- | When you want to know... | Run |
91
- | --- | --- |
92
- | Is everything still connected? | `token-harness ui` |
93
- | What should I do before a demanding task? | `token-harness optimize --task hard --profile quality` |
94
- | Is an integration actually working? | `token-harness verify` |
95
- | How much reducer saving has been measured? | `token-harness metrics --since 7d` |
96
- | Are safer provider updates available? | `token-harness update` |
97
-
98
- If a command finishes with **no action required**, stop there and use your coding agent
99
- normally. Token Harness should not send you around a `ui → optimize → ui` loop.
100
-
101
- ## Ask your AI to install Token Harness
102
-
103
- You can give this prompt to Claude Code or Codex:
104
-
105
- ```text
106
- Install the latest stable Token Harness from npm on this computer, then run
107
- `token-harness setup`. Do not install or replace Claude Code, Codex, or any
108
- optimization provider unless Token Harness's supported plan explicitly requires it.
109
-
110
- Explain the setup result in plain language: what was detected, what already works,
111
- what would change, and the single next step. Do not expose credentials, cookies,
112
- tokens, raw home paths, or private project contents. If setup proposes a supported
113
- configuration change, show me the short plan and ask before running
114
- `token-harness setup --yes`. After an approved change, verify it and open
115
- `token-harness ui` once. Then tell me clearly that setup is complete and that I should
116
- continue using my normal coding-agent command. Do not invent additional Token Harness
117
- steps when no action is required.
103
+ ```sh
104
+ git clone https://github.com/giuliastro/token-harness.git
105
+ cd token-harness
106
+ npx --yes pnpm@10.33.4 install --frozen-lockfile
107
+ npm start
118
108
  ```
119
109
 
120
- The AI should ask before the `--yes` step because that is the point where coding-agent
121
- configuration may change.
110
+ This uses the clone, not an older global installation. An unmerged branch or unpublished main
111
+ change is not automatically available through `token-harness@latest`.
112
+
113
+ ### Advanced and AI-assisted use
114
+
115
+ The long CLI flag combinations are **not** the normal human interface. They are the controller API
116
+ used by the app, automation, and optionally by a coding harness. Humans can keep using the browser
117
+ and the two entry points above.
118
+
119
+ For in-session use, this repository now includes a portable Agent Skill at
120
+ [`skills/token-harness/SKILL.md`](skills/token-harness/SKILL.md). A compatible Claude Code or Codex
121
+ skill mechanism can load it on demand, after which you can simply ask the harness to **use Token
122
+ Harness for this task**. The skill classifies substantial work conservatively, calls the existing
123
+ local `--json` optimizer at meaningful task boundaries, and can use explicit workload scheduling
124
+ when you have actually supplied a backlog. It does not run Token Harness before every tool call.
125
+
126
+ The skill is deliberately thin: Token Harness remains the deterministic policy engine. No MCP
127
+ server, background model, persistent agent, or second quota formula is added just to make this work.
128
+ Local tokens are still not subscription quota, and raw Claude/Codex percentages are still not a
129
+ common currency.
130
+
131
+ Agent use is read-only by default. If Token Harness recommends a persistent model, reasoning, or
132
+ verbosity change, the harness must first build a reviewed plan, explain the exact proposed mutation,
133
+ and wait for your explicit approval before `apply`. A saved preference may affect future sessions;
134
+ it is not silently presented as a live change to the current session.
135
+
136
+ The guided app can preview **Enable in-session guidance** for each detected Claude Code or
137
+ Codex installation. It installs the same portable skill into the agent's documented user-level
138
+ Agent Skills directory through the normal transactional plan/apply path. The agent card then shows
139
+ the live guidance state: **Enabled** when the exact skill is currently owned by Token Harness,
140
+ **Enabled externally** for a byte-identical user-owned skill, or an explicit not-enabled,
141
+ custom/conflict, or unavailable state. Existing `token-harness` skill directories are never
142
+ overwritten or silently adopted, and matching bytes alone never create an ownership claim. The
143
+ browser remains fully usable without any skill or second AI subscription. See
144
+ [RFC 0023](docs/rfcs/0023-guided-agent-skill-install.md) for the install and ownership boundary.
145
+
146
+ The older automation contracts remain available: `setup`, `optimize`, `plan`, `apply`,
147
+ `verify`, `metrics`, `rollback`, and their JSON reports. `ui --json` preserves its existing
148
+ schema-1 report; `ui --read-only` opens the legacy read-only dashboard. `ui --no-open` starts
149
+ the guided app without launching a browser. Stop either local server with Ctrl+C. See
150
+ [RFC 0022](docs/rfcs/0022-agent-native-skill.md) for the agent-facing safety boundary.
122
151
 
123
152
  ## What normal output looks like
124
153
 
@@ -167,26 +196,10 @@ token-harness ui --json
167
196
 
168
197
  `--json` keeps the complete schema-1 result and diagnostics; it is not shortened.
169
198
 
170
- ## Three commands to remember
171
-
172
- ```sh
173
- token-harness setup
174
- token-harness ui
175
- token-harness optimize
176
- ```
177
-
178
- | Command | Answer |
179
- | --- | --- |
180
- | `setup` | Is Token Harness ready, and what is my one next step? |
181
- | `ui` | What is active, how much allowance is visible, and do I need to do anything? |
182
- | `optimize` | What is the best evidence-based action for the task I am starting? |
199
+ ## Two entry points to remember
183
200
 
184
- For example:
185
-
186
- ```sh
187
- token-harness optimize --task hard --profile quality
188
- token-harness optimize --task mechanical --profile economy
189
- ```
201
+ `token-harness` opens the application. `token-harness savings` prints recorded results.
202
+ The advanced commands below are implementation tools, not a required user workflow.
190
203
 
191
204
  ## Safety and privacy
192
205
 
@@ -194,19 +207,23 @@ Token Harness is conservative by design:
194
207
 
195
208
  - normal read-only commands do not change coding-agent or project configuration;
196
209
  - `setup --yes`, `apply --yes`, `update --yes`, `rollback --yes`, and
197
- `uninstall --yes` are the explicit configuration-changing forms;
210
+ `uninstall --yes` are the explicit CLI configuration-changing forms; the guided UI uses
211
+ a reviewed preview and explicit **Approve and apply** instead;
198
212
  - plans are checked again immediately before they are applied;
199
213
  - existing files are backed up before a managed write;
200
214
  - only exact Token Harness-owned entries are removed by `uninstall`;
201
215
  - newer or untested combinations are reported, not guessed;
202
216
  - an available provider update outside reviewed compatibility is kept out rather than
203
217
  forced, and the installed working version stays in place;
204
- - the dashboard binds only to the local loopback address and provides no mutation API;
218
+ - the guided app binds only to 127.0.0.1 and protects its fixed local controls with exact
219
+ Host/Origin checks, a per-process anti-forgery token and single-use approval tickets;
220
+ - the legacy read-only dashboard and external status seam remain read-only;
205
221
  - source code, prompts, command contents, credentials, and cookies are not sent to a
206
222
  Token Harness service.
207
223
 
208
224
  Plans, receipts, metrics, and backups stay in the local Token Harness state directory.
209
- See [RFC 0004](docs/rfcs/0004-safety-and-installation.md) for the execution model and
225
+ See [RFC 0013](docs/rfcs/0013-guided-local-experience.md) for the local browser trust boundary,
226
+ [RFC 0004](docs/rfcs/0004-safety-and-installation.md) for the execution model and
210
227
  [RFC 0006](docs/rfcs/0006-cli-contract.md) for CLI/JSON guarantees.
211
228
 
212
229
  ## Supported optimizations
@@ -250,8 +267,111 @@ Most people do not need this section. Run `token-harness <command> --help` for d
250
267
  | `handoff` | Build a bounded cross-harness handoff | No |
251
268
  | `benchmark*`, `transfer*` | Capture and compare empirical evidence | Local state only |
252
269
 
270
+ ## Workload-aware allowance planning
271
+
272
+ If you know how many accepted tasks remain, the advanced CLI can ask whether that backlog fits the
273
+ **currently observed** included allowance:
274
+
275
+ ```sh
276
+ token-harness optimize --harness codex --task standard --tasks-left 5
277
+ token-harness schedule --current codex --candidate claude --task-class standard --tasks-left 5
278
+ ```
279
+
280
+ `--tasks-left` is explicit workload intent for a backlog of one task class. Token Harness does not
281
+ infer it from `ccusage`, local tokens, session length, or raw provider percentages. A workload-driven
282
+ recommendation requires complete project-local benchmark evidence for the exact model + reasoning
283
+ effort + verbosity policy in both the five-hour and weekly windows. If that evidence is incomplete,
284
+ capacity stays unknown.
285
+
286
+ When evidence proves that the current policy cannot cover the stated backlog, `optimize` protects
287
+ capacity instead of spending a quota-derived effort bonus and reports whether the five-hour,
288
+ weekly, or both windows are limiting. `schedule` can use the same target to consider the other
289
+ harness, but only when that candidate has enough conservative accepted-task capacity and passes the
290
+ existing quality, pace, availability, and transfer checks.
291
+
292
+ ### Mixed task-class backlog
293
+
294
+ For queued **new tasks** spanning more than one class, give `schedule` the mix explicitly:
295
+
296
+ ```sh
297
+ token-harness schedule --current codex --candidate claude \
298
+ --workload mechanical=2,standard=3,hard=1
299
+ ```
300
+
301
+ Mixed mode does not sum per-class task capacities as if they were separate quota buckets. It charges
302
+ each proposed task's project-local p75 cost against the same shared five-hour and weekly allowance
303
+ of that harness, then returns `stay`, `split`, `switch`, `shortfall`, or
304
+ `insufficient-evidence`. Candidate assignments require at least three coherent quality-gated
305
+ observations for the exact task class plus complete five-hour and weekly capacity evidence.
306
+ Unproven work remains visibly unallocated.
307
+
308
+ This mode is for queued/new tasks, not an in-progress handoff. Therefore `--workload` is mutually
309
+ exclusive with `--task-class`, `--tasks-left`, manual pace/quality flags, and handoff/transfer
310
+ flags. The allocator is deterministic and conservative; it does not claim globally optimal routing,
311
+ launch either harness, or compare raw Claude and Codex percentages.
312
+
313
+ No capacity after a future reset is assumed. Re-run the observation after the reset rather than
314
+ treating a forecast as provider quota. See [RFC 0020](docs/rfcs/0020-workload-aware-allowance.md)
315
+ and [RFC 0021](docs/rfcs/0021-mixed-workload-allocation.md).
316
+
317
+ ## Applying native recommendations
318
+
319
+ `optimize` remains read-only. Review a plan before applying a supported native change:
320
+
321
+ ```sh
322
+ token-harness plan --harness claude --native-policy --task mechanical --profile economy
323
+ token-harness apply --plan <printed-plan-id> --yes
324
+ ```
325
+
326
+ `apply --plan <id>` restores the reviewed harness/provider selection automatically; you
327
+ should not have to repeat `--harness`, `--provider`, `--native-policy`, `--task` or `--profile`.
328
+ Run it from the same project as `plan`. Conflicting explicit selectors are rejected, and
329
+ actual version, ownership and configuration changes still invalidate the plan. Existing
330
+ schema-1 plans remain usable; only their approved actions can execute.
331
+
332
+ The first Claude path supports the **persisted user effort preference** on the reviewed
333
+ Claude Code 2.1.261 build. It does not change model, authentication, hooks, endpoint or billing.
334
+ `max` is never persisted. Project/local/ancestor settings, custom configuration roots and
335
+ known environment/thinking overrides block the change rather than being overwritten. The
336
+ preference affects future sessions unless overridden: reopen Claude and check `/effort`.
337
+ This is not evidence of a running session's effective effort or a guaranteed quota saving.
338
+
339
+ For Codex, the same plan/apply flow manages the existing reviewed reasoning-effort and
340
+ verbosity fields through native `config/batchWrite`; project/profile overrides remain yours.
341
+ `rollback --yes` restores the complete pre-change files. `uninstall --yes` removes only owned
342
+ changes and restores a prior Claude effort preference without undoing unrelated later edits.
343
+
253
344
  ## Troubleshooting
254
345
 
346
+ ### Claude allowance is unavailable
347
+
348
+ The dashboard now explains whether the optional companion is missing, lacks the safe CLI
349
+ flags, cannot find Python, has no usable Claude session, reports an expired session, or returns
350
+ an unsupported source. It does not expose credentials, raw companion errors or private paths.
351
+
352
+ As observed on **September 5, 2026**, npm `cclimits@1.7.0` includes the merged Claude
353
+ zero-configuration support and the read-only flags. The latest GitHub Release listing is older
354
+ and is not evidence of what npm ships. To check the same path Token Harness uses:
355
+
356
+ ```sh
357
+ npm list --global cclimits
358
+ cclimits --claude --json --no-cache-write --no-stale-fallback
359
+ token-harness budget --harness claude --verbose
360
+ ```
361
+
362
+ An explicit optional installation/update is `npm install --global cclimits@1.7.0`.
363
+ Token Harness does not install it automatically or retry without its read-only flags.
364
+ A fresh local Claude cache is shown as **cached**, never promoted to live quota pacing.
365
+ A missing observation is not zero remaining allowance. Never paste credentials to debug it.
366
+
367
+ ### Codex is configured but its hook does not run
368
+
369
+ `token-harness verify --harness codex --verbose` now reads native `hooks/list` where the
370
+ installed app-server exposes it. Disabled, untrusted and modified hooks are distinguished from
371
+ an unavailable observation. Trust must still be granted explicitly in Codex. Enabled/trusted
372
+ metadata does not prove interception, reduction, or task quality; the integration remains
373
+ `config-only` until attributable runtime evidence exists.
374
+
255
375
  ### `token-harness` is not found
256
376
 
257
377
  Check that Node is new enough and the package is installed:
@@ -339,3 +459,22 @@ behavior or architecture.
339
459
 
340
460
  [Apache License 2.0](LICENSE). Referenced provider tools are independent projects with
341
461
  their own licenses.
462
+
463
+ ### Loading, impact and sharing
464
+
465
+ The dashboard shows animated, named checks while it reads your setup. Agent cards and saved
466
+ reduction records appear as they are ready; a slow allowance check does not hide the results.
467
+ Refreshing keeps previous readings visible until newer ones arrive. Errors and waiting are
468
+ explicit, and reduced-motion preferences disable animation without removing status text.
469
+
470
+ A result can now say **"65% less tool output"**, with its source, before/after values and count
471
+ of recorded changed outputs immediately beside it. An estimate says **"Estimated"**. This
472
+ percentage describes only those recorded outputs, not your whole coding session, subscription
473
+ allowance or money. Provider rows remain separate; negative results and errors remain visible.
474
+
475
+ Choose **Share result** to preview the exact summary and a locally generated image. **Open X
476
+ draft** prepares a short post. **Open Reddit** prepares a title/link; copy the summary into a
477
+ text post and choose a community yourself. **Copy for Discord** prepares a message to paste
478
+ in your chosen channel. **Save image** creates a PNG you can attach yourself. Nothing is
479
+ posted or uploaded automatically, and sharing excludes private paths, code, prompts and
480
+ account/allowance information. An open share preview stays fixed even if readings update.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "token-harness",
3
- "version": "0.1.8",
3
+ "version": "0.1.10",
4
4
  "description": "Quota-aware efficiency layer for Claude Code and Codex subscription limits.",
5
5
  "license": "Apache-2.0",
6
6
  "type": "module",
package/sbom.json CHANGED
@@ -1,14 +1,14 @@
1
1
  {
2
2
  "bomFormat": "CycloneDX",
3
3
  "specVersion": "1.5",
4
- "serialNumber": "urn:uuid:dfdf5e8b-aba7-10dd-1c9d-dd9a38f31fb3",
4
+ "serialNumber": "urn:uuid:9290051f-d6be-441a-20e1-add1f9ea0ff8",
5
5
  "version": 1,
6
6
  "metadata": {
7
7
  "component": {
8
8
  "type": "application",
9
9
  "bom-ref": "token-harness",
10
10
  "name": "token-harness",
11
- "version": "0.1.8",
11
+ "version": "0.1.10",
12
12
  "description": "Quota-aware efficiency layer for Claude Code and Codex subscription limits.",
13
13
  "licenses": [
14
14
  {
@@ -20,7 +20,7 @@
20
20
  "hashes": [
21
21
  {
22
22
  "alg": "SHA-256",
23
- "content": "dfdf5e8baba710dd1c9ddd9a38f31fb3aded525ea28f8b15a0fc5cff743444c6"
23
+ "content": "9290051fd6be441a20e1add1f9ea0ff88e9ae931921532ed8422d86db26f1b3a"
24
24
  }
25
25
  ]
26
26
  },